Guardrails AI: Deploying Generative AI Safely with CEO Shreya Rajpal

The MAD Podcast with Matt Turck · with Shreya Rajpal, Co-founder & CEO, Guardrails AI

Shreya Rajpal is the Co-founder & CEO at Guardrails AI. We cover why 10-minute GenAI prototypes fail to deliver consistent customer value, why fine-tuning does not eliminate hallucinations, and how runtime validators can enforce organization-specific rules for accuracy, compliance, privacy, and brand safety.

Watch on YouTube

Chapters

  1. 0:00 — Full episode

Transcript

Full episode

Matt Turck [0:57] Shreya, welcome to The MAD Podcast.

Shreya Rajpal [1:01] Yeah, thank you. Thank you for inviting me. Really excited to be here.

Matt Turck [1:38] So we are going to talk about the all-important topic of safety for the deployment of generative AI in production, which feels like a particularly timely conversation as people are starting to think about how you actually deploy those systems for mission-critical applications in production at scale and all the things. So you are the co-founder and CEO of Guardrails AI, a startup that aims to provide, as the name conveniently indicates, guardrails around large language models. And I'd love to start maybe with that aha moment, for lack of a better term, when you realized you wanted to build this, particularly as it relates to your background.

Matt Turck [2:02] I think I read somewhere that you had spent time in the world of autonomous vehicles. Was that the trigger? What led you to start the company?

Shreya Rajpal [2:22] Yeah, I'm happy to dive deeper into that. So my background did play a key part in thinking of guardrails and then thinking of what practically it would end up looking like as these models and applications built on top of these models enter our day-to-day software stack. So it really started around the end of last year. Like everybody in tech, I was very excited about the coming of generative AI, and especially as it relates to removing a lot of the previous barriers with doing machine learning.

Shreya Rajpal [2:58] So for the longest time, training the model, getting the data ready, et cetera, would always end up being this key blocker before you could actually build machine learning models. And suddenly it seemed like all of those blockers were removed, which was very exciting. And as I was building my own applications, I was looking to see, okay, what is still challenging in this world, right? What would the net-new problems end up being?

Shreya Rajpal [3:20] And very quickly I came upon the idea: yeah, you can build very exciting prototypes in under 10 minutes that allow you to chat with your dataset. But that ends up being very insufficient because it doesn't actually, to an end customer or to an end user, give you value consistently, repeatedly, et cetera. And so it ends up being in this realm of, yeah, it's an exciting toy, and it's very fascinating that you can get so much leverage out of very little technological investment.

Shreya Rajpal [3:53] But in order to build lasting products, you would need to have much more reliability. You would need to be able to hold applications built on top of GenAI to the same standard that we hold software engineering products, right? We live our lives on them, and we need to get GenAI-based technology to that same level. So that problem seemed like the key missing problem where, yeah, it works, but does it work enough, and does it work reliably enough?

Shreya Rajpal [4:04] So I think that was the problem. And then in terms of solutions, I'm obviously biased because I worked in self-driving, but the way of building products on top of ML seemed more similar to self-driving than you would initially suspect, wherein it wasn't just, here's this great model and the output of the model is the commodity by itself, which is what you had in the previous generations of machine learning, where you had maybe churn prediction or something, and you do analytics on top of that churn prediction model.

Shreya Rajpal [5:00] Instead, now you had this model which was kind of this component of a subsystem that is drawing upon some other predictive components and some data, and you're chaining a bunch of these together. And so that ends up being a much, much more complex system. And in that, errors and reliability and just working with the non-determinism of machine learning ends up being way more important. So this setup was very similar to self-driving, which is, you kind of take this very vast problem of self-driving, break it up into chunks, and apply machine learning very strategically and combine it with other systems for verification, et cetera.

Shreya Rajpal [5:34] And I wanted to build a framework that allows you to do something similar for all of the amazing LLM applications that we're seeing. So that was some of the motivation. I think the initial analogy made a lot of sense, but it was definitely inspired. But it's been great to see that as the stack has started to mature a little bit, as people are really thinking about how do we put these things into production, the trajectory of issues and concerns that people are running into are similar to the path in self-driving, which is, how do I get this runtime safety?

Shreya Rajpal [6:11] How do I get runtime constraints, et cetera? How do I make this really flexible technology work with what I know to be true of the world? And making that work, I think, that analogy ends up being pretty powerful. So that was those two things: my own experience with self-driving, as well as building out these applications and figuring out where the gaps are.

Shreya Rajpal [6:28] I think that led to this aha moment of, okay, we need a framework like this if we were to build LLMs safely. Yeah.

Matt Turck [6:54] Great. Great. So I'd love to spend a little bit of time on the problem itself. So as users, we're all now very familiar with the concept of hallucination, just using ChatGPT or other systems and getting results of varying levels of quality, let's say. But what is the spectrum? What is the range of ways generative AI can go bad?

Shreya Rajpal [7:21] Yeah, so I think hallucination is the top reason that everyone thinks of, and it definitely is a problem. I work with these systems not just, obviously, as the creator of the open source and as a founder of the company, but also deeply as an engineer. And you do see the extent of the problem even when the solution is roughly right. It still is hallucinating facts I never gave it.

Shreya Rajpal [7:50] So that definitely is a problem. But outside of that, there's a lot of implicit assumptions we make out of our existing workforce, and translating those assumptions into GenAI-based systems is unsolved right now. So concretely, what some of these risks would end up looking like is compliance, for example, is a large one, where there's two failures with GenAI. One is that our existing frameworks of making AI compliant within specific settings suddenly break down.

Shreya Rajpal [8:22] So existing, for example, model risk management frameworks don't really apply when you haven't built the model yourself. You didn't curate the data that the model was trained on, and so you can't make any claims to that, right? I think the second one is that there's a lot of other, I guess, guarantees that you expect for a system to be compliant that you then have to impose on generative AI. So my favorite example that I like using is, if you're building a chatbot that works in financial services, then you want to make sure that that chatbot never misleads the customers by giving financial advice, right?

Shreya Rajpal [9:01] So that, for example, is one of the key risk areas. I think, in general, other risk areas are around brand risk. So, if you are building a chatbot that has a commercial application, then making sure you're not saying anything that, even if it's legal, might end up being embarrassing to the brand or might get you in the headlines the next day. That is something that's completely undesirable. So there's also this need that people have to kind of poke and prod these systems to get them to fail, right?

Shreya Rajpal [9:11] There's people that kind of—

Matt Turck [9:12] I would never do that.

Shreya Rajpal [9:16] Yeah.

Matt Turck [9:17] We've all tried. We've all tried.

Shreya Rajpal [9:37] Yeah, we've all tried. And it's very amusing to see how you get there. There's entire large subreddits dedicated to, okay, how do you get these systems to say the most embarrassing thing? So I think brand risk ends up being another key challenge. There's a few other ones around privacy, security, et cetera. So people are very, very concerned about PII, both PII going in, leaving the organization's system and going into these models, as well as these models accidentally referring to some data that they were trained on, which has happened to me in the past.

Shreya Rajpal [10:00] And I find it very, very funny and take screenshots, but don't really do anything with those screenshots. So I think that ends up being another concern.

Matt Turck [10:32] Yeah. And there seems to have been a number of early attempts at mitigating the problem. I've spent, like a lot of people in the space, a good amount of time with our vector database friends. And then this whole concept of RAG, that a lot of people have had a lot of hope for—do you want to maybe talk about what that is and what's working about it, what's not yet working about it?

Shreya Rajpal [11:02] Yeah, yeah, absolutely. So RAG, I think a lot of folks in your audience would know, but for folks that don't, RAG stands for retrieval-augmented generation, and it's essentially a method of building with large language models so that you're not directly questioning them. You're first sub-selecting context from your organization, from your application that's helpful, and putting that context into the prompt, and then asking the LLM to answer only using the context you provided. So as an example, let's say you have some internal documents, let's say a standard operating procedure at your company, and then you want employees to be able to ask questions from that standard operating procedure.

Shreya Rajpal [11:40] You would first figure out what the most relevant subsections of that SOP are and then append those SOP subsections to your prompt and then ask the LLM to answer using the subsections you provide. So that's, high level, what retrieval-augmented generation does. You retrieve the highlights and then you augment the prompt and generate. Yeah, just for summarization. So it is right now the primary way that most people are building with LLMs, especially I think one of the most common interfaces that people end up building are chatbots.

Shreya Rajpal [12:17] And the way to get really performant chatbots that actually work for your context are with retrieval-augmented generation. There aren't really too many alternatives right now to do this outside of maybe fine-tuning your model. And even then, retrieval-augmented generation, even with fine-tuning, might still be helpful to do on top of that. That's also—

Matt Turck [12:34] Just to jump into this, so this RAG, this fine-tuning, again, like you just did, to make this interesting to a broad group of people, what does fine-tuning actually mean? Everybody's sort of heard the term, but what does it actually mean?

Shreya Rajpal [13:03] Yeah, yeah, yeah, absolutely. So fine-tuning used to be, for the longest time, before prompt engineering entered the scene, the primary way to make machine learning models work for you. And the idea was that typically, in research, research labs or academia or other folks would end up designing architectures, which are ways of configuring deep learning neural networks together that typically work really well for certain problems. And then you take that architecture as is, start with a random initialization of that architecture, and then do a bunch of passes where you just input your data and then try to make that architecture work really, really well for your data.

Shreya Rajpal [13:39] So that was more traditional. That was what training a deep learning model meant. I think then we ended up, in the trajectory of building these models, getting to a point where the model sizes ended up just being so large and so massive. And the data that you'd need to actually get meaningful signal from those large architectures would end up being small, right? So your organization data would be much smaller than the size of the internet, for example.

Shreya Rajpal [13:50] So then we ended up in the space where we started doing a lot of pre-training. So you take these architectures, pre-train them on very, very large generic datasets, and then once they're pre-trained, you'd essentially take the small amount of custom organization data that you have, and then run a few passes with very low learning rates that allow you to just kind of sharpen the models to your specific data.

Shreya Rajpal [14:38] So this kind of paradigm of pre-trained models followed by fine-tuning that you do yourself ended up emerging as these model sizes grew pretty substantially. Yeah, so the key differences are: RAG, for example, is a runtime thing, whereas fine-tuning is an offline process. So before you're able to launch and release your application or your model, you would typically spend some time figuring out what the right data that you want to fine-tune is, and then training your model on that data, and doing a bunch of experiments to make sure that you are able to get good model performance, et cetera.

Shreya Rajpal [15:03] So the level of investments ends up being pretty different. Yeah.

Matt Turck [15:35] Great. Thank you very much for that. So if I'm an enterprise and I really want to deploy my own internal chatbot, let's say, and I have resources, I have engineers, and I do that kind of belt-and-suspenders approach of doing both RAG and fine-tuning, what does that mean in terms of hallucination? Are there benchmarks? Does that mean I have, like, 5% hallucination or 15%, or how do people measure the severity of the problem?

Shreya Rajpal [16:05] Yeah, I think it's a great question. I think you're kind of hitting on what the big challenge in the space is right now, where evaluation is this big open problem. And we're at that point where we're already beyond where traditional academic metrics were. And so it's kind of like the Wild West a little bit out here. In terms of organizations that are doing this two-step approach of first fine-tuning a model and then doing RAG on top of that, first, props to them.

Shreya Rajpal [16:37] I think it's definitely, especially for fine-tuning, a big challenge ends up being just getting the dataset right. So for the last decade or so, we've heard that data is oil and data-centric machine learning, and just getting that dataset ready and together ends up being a substantial investment. Obviously, that investment often results in that lift in model performance, and it allows you to work with much smaller models that work better for your data, that often have much lower latencies.

Shreya Rajpal [17:09] But in order to get those results, you typically need to make that initial investment. I think where it's hard to quantify, it's hard to give a blanket quantification of how much hallucination decreases by fine-tuning, because it depends on the type of task that you're doing fine-tuning on. So to make it more concrete, there's different types of hallucinations, and depending on if your task is very, very structured versus very unstructured—and I'll get into that in a second—you might end up seeing different lifts with fine-tuning.

Shreya Rajpal [17:49] So this is anecdotal, but some behaviors I've seen in my work is that if you have a very structured task that involves either generating code in a specific language or generating a specific structured output like JSON or CSV or something, that with fine-tuning ends up being pretty effective. If you have an unstructured task where you essentially still want to do question answering but only on your data, that typically requires a much bigger dataset.

Shreya Rajpal [18:25] And often it's sometimes harder to get that data curated as well, right? So depending on both the amount of data that you have for fine-tuning and the type of the task, whether it's structured generation versus unstructured text, you might end up seeing different lifts. So it's very possible and likely that you're still going to run into hallucinations even with fine-tuning. It's just going to understand your task domain a little bit better.

Shreya Rajpal [18:55] So at the end of the day, these models are—I don't personally think this is a hot take, but for some reason this ends up being a spicy take—which is that, at the end of the day, these models are like next-token predictors, which is, they kind of look at what they've predicted until now and then figure out what the next token thereon is. And from that, even with fine-tuning, that fundamental behavior remains the same, right?

Shreya Rajpal [19:23] So the hallucinations happen because of this behavior, and fine-tuning kind of doesn't change that. This is kind of like a non-technical example I like using to give an example of why this hallucination happens. Somebody I know was asking one of these AI models for kid-friendly activities or family-friendly activities in some city that they were visiting. And as this model was generating a list of outputs, essentially one of the outputs it shared was, okay, go to the zoo, go to the X Zoo.

Shreya Rajpal [19:51] And then it said, go to this city zoo with your family, and it's a good event. And then when they actually looked it up, that city never actually had that zoo. But we could have. Yeah, I think the city is lagging the AI model. Like, once the AI model has decreed there must be a zoo, the city planning should get on that. Yeah, but it's so conditioned to think about, okay, kid-friendly activities must include a trip to the zoo.

Shreya Rajpal [20:23] Fine-tuning won't make that go away, right? Fine-tuning won't make it go away that, okay, a zoo doesn't actually exist. So hallucinations aren't removed by fine-tuning, for the same reason that they aren't with RAG. Like, with RAG, you end up getting more context and it's at runtime, so it ends up being a little bit more concrete, but it doesn't completely eliminate hallucinations. I always feel a little bit nervous sharing numbers like 5% or not because it's so context-dependent.

Shreya Rajpal [20:46] But in my work, and I'm keeping this very vague because different contexts, different applications are going to see different numbers, around 30% or so hallucinations with RAG for a specific application. Yeah.

Matt Turck [21:09] And I assume the answer is probably no. Actually, I don't know. Maybe I'm wrong. But given the above, if I'm an enterprise and I'm thinking whether I want to work with GPT-4, whichever other proprietary API kind of system, versus grabbing something off Hugging Face open source and just doing my own work and/or Llama 2 or whichever, are all those models created equal in terms of hallucination? Or is it like open source that can heavily customize, fine-tune to my own needs is less likely to hallucinate?

Matt Turck [21:25] Is there a difference, or does it depend on what you do with them?

Shreya Rajpal [21:50] So, it always depends on what you do with them, but right now the models aren't equal. I think OpenAI is definitely leading the charge in terms of how effective—I guess just how flexible and generalizable—the models are. But other players, including open source, are quickly catching up. Let's see. I think there's a benchmark called LegalBench. Definitely recommend everybody goes and checks it out.

Shreya Rajpal [22:16] LegalBench was a benchmark that was released on a myriad of legal tasks, like any task that could be augmented or automated by AI in the legal domain. There's some benchmark for that in this LegalBench set. And in that, they essentially evaluated a bunch of models, including open source. And you would typically see that OpenAI-based models tend to do very well on these tasks that they haven't seen before, that are newly created benchmarks.

Shreya Rajpal [22:47] So I think OpenAI, and specifically GPT-4, is pretty powerful and potent in terms of its flexibility. In terms of open-source models, there's this one study that came out recently. It was a blog actually published by the Anyscale team. I would also definitely recommend folks go out and check that out. It was essentially a blog that looked at summarization and, if I'm remembering correctly, hallucinations in summarization and evaluating hallucinations in summarization while using GPT-4, which once again is the leading model out there, versus Llama 2 fine-tuned on some dataset.

Shreya Rajpal [23:14] If I remember correctly. Again, I'm kind of getting this off the top of my head, so please forgive any inconsistencies.

Matt Turck [23:27] This is great. Well, we'll link it in the show notes, but that's great.

Shreya Rajpal [23:40] Yeah, awesome. Yeah, 5 or 4 today. But with fine-tuning, if you make that investment in curating your dataset, in running that fine-tuning job, and then serving it, you're able to, on some of these tasks, have comparable performance. So I think that's where the strength lies, which is, if you're able to make that investment, then the technology ends up being cheaper to run long term, and you end up having more control over it.

Shreya Rajpal [23:57] So that's how I would generally think about the differences.

Matt Turck [24:09] Yeah. Okay, great.

Shreya Rajpal [24:09] All right.

Matt Turck [24:31] So, jumping into Guardrails more specifically, tell us about the general approach, and then after that I'd love to actually jump under the hood and go into—you have concepts of guards and concepts of rails. I'd love to talk about this, but let's start with a general idea of what it is that you're building.

Shreya Rajpal [24:54] Yeah, absolutely. So I think the core insight that we had was that the models are insufficient by themselves, right? The models are great and really performant, and you can fine-tune them and do a bunch of other things. But at the end of the day, what you really need is that the final outcome that you're getting from the model fits within the constraints of your systems, right? And those constraints might be accuracy or hallucination, or they might be all of these other failure modes that we talked about early on.

Shreya Rajpal [25:16] So the key insight was, yes, you can use whatever model you want, but then at the end of the day, you'd want to make sure that all of these functional risks or harms that you care about—there shouldn't be any hallucination. If I'm sharing some information to my end customer, I shouldn't ever be misrepresenting what I have or what I'm building to the customer, or my chatbot shouldn't be saying anything that's embarrassing to my company, right?

Shreya Rajpal [25:50] These are all of the areas, and you can add this to the prompt. But at the end of the day, you'd want to have independent systems that check to see if any of those things are violated and make sure that if they are, then you have a very flexible set of policies to mitigate and stop that harm from occurring, right? So that was the general idea of Guardrails, where, A, it works on a layer which surrounds the core LLM and therefore allows you to have much safer guarantees on top of the outputs you're getting from the LLMs.

Shreya Rajpal [26:24] And B is that it basically takes what is, at the end of the day, very flexible general-purpose models and makes sure that whatever custom criteria, custom rules that you have either from your organization or the domain that you're working in, you're able to take those rules and implement them and execute them in code, right? Which is, you're taking a general model of the internet and making sure that it doesn't violate any policies that are internal to your organization, perhaps.

Shreya Rajpal [26:46] So that was the general idea behind Guardrails, which is that you would need this independent validation and verification layer to make sure that you can actually use these models in production. So that's—

Matt Turck [26:47] And do it at runtime.

Shreya Rajpal [27:00] Exactly, and do it at runtime. I think the runtime part is critical because, once again, no amount of offline evaluation gives you that confidence that you need for shipping a lot of these things into production. Yeah.

Matt Turck [27:02] So how does that work?

Shreya Rajpal [27:28] Yeah, great question. So I think what the framework offers is a few different sets of tools to make this problem tractable. So the first is the concept of runtime guards. What runtime guards do is they surround your AI model and make sure that any of these functional concerns that you have are verified. So a runtime guard is typically created by combining a few different validators together.

Shreya Rajpal [27:40] And those validators typically check for each constraint.

Matt Turck [27:42] So what's a validator?

Shreya Rajpal [28:05] Yeah, so a validator is an independent check that analyzes your output and scores it for one particular risk area. So any one specific risk turns into one specific validator. So, for example, hallucination. You might have this one risk of hallucination and making sure that any output that's generated is true given the source that you gave it. So what the hallucination validator will do—we call it Provenance.

Shreya Rajpal [28:23] So what the Provenance validator does is it makes sure that there's no hallucination in the output that you're generating, right? There might be one more validator that is around, let's say, not mentioning any peer institutions. So if you're a company, you don't want somebody to ask you a question about, like, okay, who's the best tennis shoe company, and you're Nike and somebody else says Asics or something, right?

Shreya Rajpal [28:58] So that's just embarrassing. So there might be some reference of not making a reference to your peer institution. And so that would be one specific validator. Another validator might end up being just a very, very simple regex rule. So making sure that you're only allowed to answer in outputs that match a specific pattern, right? So if you're extracting some information about an interest rate, you always know that an interest rate looks something like this.

Shreya Rajpal [29:28] And so you want to make sure that you're extracting information that fits that criteria. So all of these end up being very independent checks that check for accuracy or performance or private information or hallucinations or compliance constraints, and there's a whole catalog of validators that we have in Guardrails. And you can take them, mix and match them, configure them, and combine them together into a guard that runs at runtime and makes sure that all of those criteria that you specifically want to check for, none of those criteria is violated, essentially.

Matt Turck [29:58] And double-clicking on the Provenance one, so how does that work? Like, you go search into the source of truth and you match the answer against whatever comes out of it?

Shreya Rajpal [30:19] Yeah, yeah, yeah. So Provenance is really exciting. We've done a bunch of work on Provenance, which is available in the open source. So Provenance, how it works is that any sentence or any utterance from the LLM should have some provenance that establishes, okay, this is the source where this sentence or utterance came from, right? So it's most effective in RAG, or retrieval-augmented generation systems, where you give the LLM some context and you want to make sure any output is coming only from the context and not from the general world knowledge that the LLM was trained on, right?

Shreya Rajpal [30:58] And this actually ends up being a pretty tactical problem, even in practice. So, for example, you often see that even after you give it some relevant context, if somebody asks a question, it'll have some sentences from that context, but it also has general definitions of what something means from the internet. And so it'll add those sentences in there, which ends up being misleading to your customers.

Shreya Rajpal [31:29] So how the provenance validator mitigates this risk is it looks at your output sentence by sentence, and then, using a bunch of different machine learning techniques, establishes that, with this level of confidence, your sentence came from this part of your standard operating procedure, or this section of your help center articles, or this section of your document, et cetera. So overall, that allows you, as a developer or as somebody building this application, to, A, filter out any hallucinated sentences and only give the sentences that came from something that you know to be true and you know your company can support.

Shreya Rajpal [32:04] And B, it also allows you to basically build interesting analytics on top of the thresholds. So you can say that typically I tend to find that I only have 80% confidence even in the sentences I know to be true, so should I generally end up increasing that confidence? Yeah.

Matt Turck [32:19] Okay. And you mentioned machine learning techniques. It sounds like it's a little bit of a mix of, like, some validators maybe rules-based and others machine learning-based. Is that fair?

Shreya Rajpal [32:48] Yeah, yeah, yeah. There's actually even more. So, rules-based, machine learning-based. We also sometimes hook up, when possible, when it makes sense for the validator, into some external systems. So, for example, we've done a bunch of work in text-to-SQL, and we found that our text-to-SQL validators were very, very competitive with what the state of the art is using generic off-the-shelf models. And so how text-to-SQL works is we essentially create a sandbox of your database using your SQL configuration, and then any SQL code that is generated by the LLM is first run against that sandbox.

Shreya Rajpal [33:16] Any errors are corrected, and then we allow the LLM to fix those errors using the techniques available in Guardrails. So that's getting a much bigger lift over directly using the LLMs. Yeah.

Matt Turck [33:41] Okay. So some of it is a little bit machine learning to correct machine learning. So, how do you think about this? Because the goal is to get to 100%, presumably, accuracy, but you have something that's not at 100%. You use something that's typically not 100% either to correct it. So how do you think about getting to that 100%, or can one eventually get to 100%?

Shreya Rajpal [34:06] Yeah, yeah, I think it's a great question, honestly. And it's also the key limitation of working with machine learning, which is 100% is almost never achievable. And so in practice, what you end up doing is, A, trying to, with an individual model, improve performance as much as possible by training on better data, by using better techniques, and doing a bunch of other things. And B, and this is one of the things that Guardrails really hooks into, is using this machine learning concept of ensembling models together.

Shreya Rajpal [34:38] And what that typically—I think my favorite way of explaining what that does—is that you can think of ensembling as stacking a bunch of different sieves on top of each other, where each sieve has holes in different areas, right? So this stack of different sieves ends up being so that your holes are all non-overlapping. And so you end up getting something that is greater than the sum of its parts and ends up being more watertight and more correct.

Shreya Rajpal [35:10] So ensembling kind of works in a different way, which is you have your large language models such as OpenAI or Anthropic or Cohere or Llama 2, but that has specific failure modes. You combine it with these guardrails and validators that also under the hood use machine learning, but they have maybe some other different failure modes. And when you run this whole ensemble together, you end up getting something that is much more performant and tends to have much, much higher robustness compared to just using the model by itself.

Shreya Rajpal [35:41] Yeah, so I think in general this idea resonates a lot with this recent blog and writing from the Hazy Research Lab at Stanford, which I also recommend people check out, which is LLMs and generative AI really help solve the first-mile problem in ML, right? But the last-mile problem, which is like, how do you take this generic, generalizable technology and make it work specifically for your use cases and for your actual application, you need other more traditional ML tools to make sure that you have the guarantees and the reliability that you typically need.

Shreya Rajpal [36:02] Yeah.

Matt Turck [36:33] How does Guardrails manifest? We talked about failure modes earlier, but sort of generically across RAG and other things. Like, how does that manifest if the chatbot, for example, mentions the name of ASICS in the shoe example? What happens? Like, do I control this, or does it just kill the answer? Like, how does it fail? What's my sort of end-user experience?

Shreya Rajpal [36:57] Yeah. Yeah. I think that's a great question. We thought about this a bunch. So the open source supports a bunch of different policies on failure. So in the extreme, what we support is this re-asking strategy that Guardrails kind of supports really, really well, where it hooks into the LLM's ability to self-heal or correct their own outputs if you give them enough context, right? So how that works is that let's say you have an LLM output and it fails the validator that you have set up.

Shreya Rajpal [37:27] That says you can't talk about your competitor. And so what Guardrails would do, or Guardrails re-asking would do, is it would automatically construct a new prompt and only give it the information it needs to be able to correct its output. And then you send your prompt back to your LLM, get a new output, and more often than not, that output ends up being correct. So once again, it's not 100%; it doesn't work 100% of the time, but it does end up working pretty substantially in a lot of outcomes.

Shreya Rajpal [37:55] So that is one of the strategies of how do you handle failures. But there's a bunch of other things available, including programmatic fixes when possible, right? So if there's any hallucinated text that is detected, you just filter out that specific text and don't throw out the rest of the answer. In some other cases, you can alert a human and then bring a human to review that output and make sure that only when a human okays it do you send that output forward.

Shreya Rajpal [38:24] And there's a bunch of other filters around responding with canned responses, or falling on other fallback systems, raising an exception, et cetera. So all of those options are completely configurable, and you can set them on a validator level. For example, it might end up being that if you're mentioning a competitor, that is dangerous, where you'd much rather make that extra trip to the LLM rather than sending that response out.

Shreya Rajpal [39:03] And so you can configure that policy just for the competitor-check validator. A bunch of other things, like always ending a chatbot response with thank you or please, right? Maybe it's bad, but it's not as bad that you'd much rather answer the question faster rather than make that extra trip. So you might maybe programmatically inject that at the end of the sentence. And so, for each validator, you can configure what the policy needs to be.

Matt Turck [39:19] Yeah. As all of this happens at runtime, how should one, or how do you all think about performance latency? Does it have any impact?

Shreya Rajpal [39:40] Yeah, yeah, yeah. I think there's basically no free lunch in terms of safety. So typically, running these validators ends up incurring some additional cost, both in terms of latency as well as compute, or making another call to the LLM, et cetera. So we generally think about it in two ways, which is, one, a lot of the validators are only useful when you have these specific runtime constraints, right?

Shreya Rajpal [40:07] And we've done a bunch of optimizations to make it as fast as possible, including being able to parallelize all of these validators so that you end up only getting the minimum runtime that you need, so that you end up getting something that works within your constraints, right? So, for example, for hallucinations, we essentially have multiple implementations of those hallucination detectors, or the provenance validators, that allow you to say, okay, if I want the really, really great-performing outcome, I might need that extra cost and that extra time.

Shreya Rajpal [40:52] But if I'm okay with, if I don't have as strict constraints around hallucinations, I can do with the really fast but maybe not that good performance. So this is also very in line with machine learning. And in machine learning, the best models are typically really large and end up having that extra latency. So I think we get some of those similar principles into Guardrails as well.

Matt Turck [41:24] Yeah. Great. So you're early in the journey of building the company, at the same time very active and shipping. Version 0 just a few days ago, so fantastic. What's the role of open source for this business? And the question behind the question is, open source is a great way to build community and users and all those things. Occasionally, open source, or maybe what it was originally, was like you had a bunch of people that were helping to build parts of this.

Matt Turck [41:50] So is there a very large number of validators, and part of the idea is that the community's going to build some of those validators? How do you think about community, not just from a go-to-market perspective, but from a product-development perspective?

Shreya Rajpal [42:19] Yeah, yeah, I think that's a great question. I think the community is definitely contributing a number of these validators. I'm not sure when the episode will go out, but this morning, we had a contribution for a regex check, making sure any output that you have matches a specific pattern that you predefine, which was contributed by the community, et cetera. So people—it's very interesting because the applications are so wide and varied, and then based on that, people end up making specific types of contributions.

Shreya Rajpal [42:54] So I think we're definitely focused on that. I also do think that what the open source really ends up doing as a framework is taking down this very abstract problem of what it means to do safe AI development, right? It's a very abstract problem. It's almost an academic problem to some degree. And it takes that and breaks that down into an engineering problem, really, right? Like, okay, if these are the risk areas that you care about, it ends up having these safety measures that are grounded in code rather than just in policy.

Shreya Rajpal [43:20] And so that, I feel like, is a big value of the open source. And that's also why a lot of people use the open source, in addition to the fact that it ends up providing them much higher performance and safety and everything. Yeah, that's really cool. So I think that allows us to kind of shape how safe AI development is done in industry, right?

Shreya Rajpal [43:44] So it allows us to basically build these systems that are much more reliable using techniques or frameworks—I guess, using the architecture stack that Guardrails kind of lays out. So that ends up being a lot of the value. Yeah.

Matt Turck [44:12] How big is safe AI as an industry? And maybe it'd be interesting for people here that want to learn if you could mention who else you think is doing interesting work. And I'm not necessarily asking to mention competitors or anything, but researchers, papers, or people doing great work in big tech, but in that sort of AI safety space.

Shreya Rajpal [44:40] Yeah, I think safe and responsible AI has always existed as an industry, but obviously, with the explosion of generative AI that we're seeing and the proliferation of these applications, it's becoming much more relevant. So there's a lot of interesting work from a bunch of different organizations. I do think the OGs in the space were Microsoft's Responsible AI team. I think coincidentally, they're also called the RAIL team.

Shreya Rajpal [45:06] No affiliation between our Rails and Microsoft's RAILS, but I think they're some of the teams that pioneered a lot of the really amazing work that Microsoft is doing, leading the charge in generative AI and making sure that all of those products are safely and responsibly developed. So I think that is pretty exciting to me. I think in research, there's a lot of exciting work, both from the security side as well as from the AI side.

Shreya Rajpal [45:34] So I definitely recommend Percy Liang at Stanford. His work, his team, for example, does a bunch of work in verifiability, which is pretty exciting. There's also, from the security side, some really exciting work by this Professor Leon Derczynski. I'm probably going to mess up the name pronunciation, but he's doing really fantastic work from the security side about how to safeguard LLMs.

Shreya Rajpal [45:54] He's a professor at the University of Washington. So I think those are some of the academic side which are pretty promising and exciting to me. Yeah.

Matt Turck [46:22] Okay, very cool. Thank you so much for sharing this. So maybe to close, taking a step back and still, as part of a general desire to make this educational for folks, you're navigating in all the right crowds. As I was preparing for this, you were partnering with LangChain, you were doing a Hugging Face meetup, you're partnering with Weights & Biases. We actually had Lukasz on this podcast a few weeks ago.

Matt Turck [46:49] Who do you see doing interesting work in generative AI these days, maybe beyond those names, which are some of the obvious ones? Whether those are friends or colleagues or people that you follow from afar that people should look up or research, anything that comes to mind, big company, small company.

Shreya Rajpal [47:09] Yeah, I think that's a fantastic question. I think we're also at a point in time where it honestly feels like a fire hose, and it's like you're keeping up with this fire hose and just getting swamped every day. So, let's see. I think on the company side, I do think LlamaIndex is really, really cool. I think, like I said, RAG is the way to build these models today.

Shreya Rajpal [47:22] And LlamaIndex is again taking a very grounded approach to doing that. So I think they're a pretty cool company.

Matt Turck [47:36] Yes. And we had Jerry on this podcast also a few weeks ago. If people are interested, you can go back a few episodes and check out our great conversation with Jerry, the founder of LlamaIndex.

Shreya Rajpal [47:59] Yeah, yeah. I think Jerry's fantastic. Obviously, I think you mentioned a lot of other fantastic people. I think LangChain is obviously leading the charge. I do think LangSmith is great. They're making a great push around LangChain, and I think that's awesome. Weights & Biases, I've long been a fan of Weights & Biases, and I think they know that as well. I've been using it for many years in addition to my ML work over all these past years.

Shreya Rajpal [48:25] I think specifically in generative AI, all of the foundation model companies are still—it's very hard to get away from the fantastic work that they're doing. So, OpenAI, obviously, like I said, their models are really competitive. But Anthropic, with their constitutional AI, I think that was a really fantastic approach to thinking about AI safety in a very principled manner.

Shreya Rajpal [48:50] And I think Cohere is doing a really fantastic job, taking a much more enterprise-focused approach to thinking about generative AI. So they recently had this model release called Coral, I think, which is a pretty fantastic implementation. I think their re-ranking model is something that I know a bunch of people I've spoken with use in their actual systems.

Shreya Rajpal [49:23] I think some other academic work that is really exciting to me is how do you take what is now the limitation of generative AI, and what's basically going to come next on the horizon, right, which allows you to train models with much longer context and allows you to iterate on the inference-time speeds of these models. I think that's pretty awesome.

Shreya Rajpal [49:40] Also, Albert Gu at CMU and his work on S4 models and state space models is, I think, really promising. So we definitely look at what's coming next on the horizon in terms of what's next beyond Transformers, essentially.

Matt Turck [49:55] Yeah. Okay, fantastic. Where can people find you online?

Shreya Rajpal [50:21] Twitter is—I'm fairly active on Twitter. You can follow my personal account, which is Shreya R, or the Guardrails account. I think we both talk about the exact same thing, so either works. I'm on LinkedIn with my name, Shreya Rajpal. And then also check out the Guardrails website and GitHub, or if you just search for Guardrails AI, you'll be the top result.

Matt Turck [50:30] Yeah. Very cool. And presumably you're recruiting. Who are you looking for?

Shreya Rajpal [50:43] Yeah, yeah, we're actively recruiting right now. We're also looking for folks on the open-source side, so people that have experience working with open-source communities.

Matt Turck [51:00] Great. And you can add /madpodcast so we get a 25% placement fee.

Shreya Rajpal [51:03] Yeah, absolutely. Yeah. All right.

Matt Turck [51:25] This has been a wonderful conversation, Shreya. Really appreciate it. And this is super interesting. I'm very excited to see how you all do in the next few years because it sounds like you're focusing on a super important problem with a really interesting approach. So I'm excited for the next steps. Thank you so much for coming on The MAD Podcast and telling us all about it. Thank you.

Shreya Rajpal [51:29] Thank you so much. Thanks again for inviting me. I really enjoyed our conversation.

Matt Turck [51:30] Yeah.

Shreya Rajpal [51:30] Bye.

Matt Turck [51:31] Yeah, bye.

Shreya Rajpal [51:36] Thanks for joining us for The MAD Podcast.

Matt Turck [51:54] We're back here every Wednesday with new conversations with leaders in the machine learning, AI, and data space. And if you like this show, you can also find a video recording of not only this episode, but many, many more over on the Data Driven NYC YouTube channel.

Shreya Rajpal [51:57] Thanks again, and catch you next week.