Top AI Researcher on GPT 4.5, DeepSeek and Agentic RAG | Douwe Kiela, CEO, Contextual AI
The MAD Podcast with Matt Turck · with Douwe Kiela, CEO, Contextual AI
Douwe Kiela is the CEO at Contextual AI. We cover why DeepSeek is an existence proof that synthetic data can produce near-frontier models, why RAG, fine-tuning, and long context solve complementary problems, and why enterprise AI needs audit trails and uncertainty flags rather than 100% accuracy.
Chapters
- 1:57 — Thoughts on the latest AI models: GPT-4.5, Sonnet 3.7, Grok 3
- 4:50 — The test time compute paradigm shift
- 6:47 — Unsupervised learning vs reasoning: a false dichotomy
- 7:30 — The significance of DeepSeek
- 10:29 — USA vs. China: is the AI war overblown?
- 12:19 — Controlling AI hallucinations at the model level
- 13:51 — RAG: definition and origin story
- 18:46 — Why the Transformers paper initially felt underwhelming
- 20:41 — The core architecture of RAG
- 26:06 — RAG vs. fine-tuning vs. long context windows
- 30:53 — RAG 2.0: Thinking in systems and not models
- 31:28 — Data extraction and data curation for RAG
- 35:59 — Contextual Language Models (CLMs)
- 38:04 — Finetuning and alignment techniques: GRIT, KTO, LENS
- 40:40 — Agentic RAG
- 41:36 — General vs. specialized RAG agents
- 44:35 — Synthetic data in AI
- 45:51 — Deploying AI in the enterprise
- 48:07 — How tolerant are enterprises to AI hallucinations?
- 49:35 — The future of Contextual AI
Transcript
Thoughts on the latest AI models: GPT-4.5, Sonnet 3.7, Grok 3
Matt Turck [1:46] Hey, Douwe, welcome.
Douwe Kiela [1:48] Hi, thanks for having me.
Matt Turck [2:18] So I thought a fun way to start the conversation would be to riff on the crazy pace of releases over the last couple of weeks. GPT-4.5 was released, so what was formerly known as Orion and now the largest OpenAI model that was released, even though they don't disclose the number of parameters. Claude 3.7 Sonnet and then Grok 3, and then going all the way back to January of this year, Deep Research, Operator, all those things. So I'm curious, maybe with your AI researcher hat on, what do you make of all of this?
Matt Turck [2:31] What catches your imagination? What do you think is more or less interesting?
Douwe Kiela [2:58] Yeah, it's exciting times, right? There's so much happening, it's hard to keep up. But yeah, I think the new model releases are interesting. I think some people are maybe a little bit underwhelmed. I think expectations are also getting maybe a little bit inflated with all of these releases. For me, the most exciting thing by far is DeepSeek, where that really, I think, changed the narrative in the AI ecosystem around what's possible and who actually is an incumbent and what is the moat that some of these companies have.
Douwe Kiela [3:32] So yeah, I think it's really great for the world that we have kind of an existence proof now that it's actually not that hard to do this. And so you don't need to invest all that much in data, and you can use synthetic data and get a pretty good model out of that. So that's really exciting.
Matt Turck [3:46] GPT-4.5 does indeed seem to be a little mixed. Claude 3.7 Sonnet, Grok 3, there's been, like, that whole discussion around scaling laws and whether we were sort of hitting a wall there. What do you think?
Douwe Kiela [4:46] Yeah, I think it depends a little bit on how you measure things, right? So I think there's this unhealthy obsession with benchmarks in the field where it's like, oh, this one didn't get much better at math or coding, but maybe it got better at all these other things that we are not actually measuring with public benchmarks, right? So I'm sure that on their internal leaderboards, those new models are a lot better than the old models. And that's why they ship them. But it might not be borne out as obviously on the benchmarks that everybody's looking at.
The test time compute paradigm shift
Matt Turck [5:13] What about all the reasoning stuff? And in particular, this whole evolution of the sort of dominant paradigm towards test-time compute and all the things that may be—defined for any listener that may not be following this sort of every five minutes on Twitter—what is test-time computing?
Douwe Kiela [5:39] So we used to all be obsessed with training scaling laws and thinking about how much data, sort of data scaling laws and parameter scaling laws. So the bigger the model and the more data you had, the better the model would get. But there are obvious limitations to that. And with o1, and I think also some other work done by other folks, it became very clear that there also is a test-time scaling law. And I think this is something that in machine learning literature has been known for a while.
Douwe Kiela [6:12] I don't think it's been labeled as a scaling law necessarily, but obviously you can also do things during test time. So when you're actually doing inference with the model, that will help you get to a better answer. And so showing that that was possible and that that would actually scale with the more time you spend thinking about something, the better your answer gets. Yeah, that's really an exciting paradigm shift, I think. So there are these trade-offs now that you have to make around train compute, train data, test-time compute, maybe test-time data as well, if you want to sort of augment the context using retrieval, maybe.
Douwe Kiela [6:31] So there's a lot of exciting new opportunities coming with this test-time compute paradigm.
Unsupervised learning vs reasoning: a false dichotomy
Matt Turck [6:52] GPT-4.5 was going to be the last model to be just focused on unsupervised learning and not include reasoning. Do you think that's where the world is going? All models are going to be a combination of unsupervised learning and sort of test-time compute kind of paradigm?
Douwe Kiela [7:23] It's not really a real dichotomy there. I think you could argue that GPT-4o is also already a reasoning model. It just hasn't been trained on reasoning specifically, but it can do chain of thought, right? So if it can do chain of thought, it's basically already a reasoning model. It just hasn't really been optimized for that. So yeah, I don't think that that's a huge shift in any way. It's just a sort of natural next step.
The significance of DeepSeek
Matt Turck [7:53] All right, so going back to DeepSeek that you mentioned a minute ago, what's sort of fascinating is that a couple of weeks ago, whenever DeepSeek came out, it felt like a Sputnik moment. Marc Andreessen said that, others said that, and the entire world freaked out for like two to three days. And fast forward to today, it almost feels like some of it, at least in the press, was swept under the rug, and we're just proceeding as originally planned. But it's still a big deal.
Matt Turck [7:59] So maybe explain why it's a big deal, and where do you see the impact?
Douwe Kiela [8:29] Yeah, I think why it's a big deal is—the reason for that is very different from what you would see in the press. I don't know why journalists do this, but they really like a very clear adversarial story. It's like the U.S. versus China, right? So something like that is very easy for people to understand. And I guess that's what people enjoy reading or what people click on, and that's why that's become part of the dominant narrative. And calling it a Sputnik moment also kind of fuels, goes in that direction of a cold war on AI or things like that.
Douwe Kiela [9:06] And I think that's all completely overblown. But I think it is very interesting technology. And so what I was saying earlier, right, it's really an existence proof. It's like, when you have a bunch of GPUs and you know how to train a language model, and if you can train up a pretty good base model, so that's their DeepSeek V3, then you can give that reasoning capabilities relatively cheaply. And by that, I don't mean in terms of compute; I mean in terms of data.
Douwe Kiela [9:41] So the data that they got came from other language models. So it's synthetic data. They also had some human annotators. They're a little bit unclear about where the data really came from. There are some questions about that in the broader community. But it just shows that you can train an amazing model on relatively little data and have that thing be almost as good, and sometimes even better, than real frontier models. So that's very exciting. But the other wrong narrative that I see everywhere is like, oh, it only costs like five-point-something million to train.
Douwe Kiela [10:17] It's like, that was a single training run. So how many training runs did they have to do to figure out what the optimal training run would be? So I would guess that they spent at least 100x the amount of that single training run, right? So it's a lot more expensive. You have to pay for people, data, all the operations, then for the data center itself. And then you need to figure out how to train the model optimally. And then you do one single training run at the end of it.
Douwe Kiela [10:28] So that single training run costing like $6 million—that's pretty expensive for a single training run.
USA vs. China: is the AI war overblown?
Matt Turck [10:44] So what's your take then on the sort of geopolitical war, or maybe lack thereof, between China and the U.S.? Is that overblown as well? Is it more of a media creation? What's your take on it?
Douwe Kiela [11:07] I think it's probably good for the world. I think the Chinese economy is very good at producing things cheaply and in a highly optimized way. And so, if you're not in the foundation model race, then it could be a very good thing. So for me, it's great if language models get commoditized, because then we can use them for all kinds of interesting applications in a much better way. So what we do is we contextualize the language models so that they can do their job, right?
Douwe Kiela [11:37] So I think that's actually where all the really interesting problems are right now. It's not even really about language models anymore. That has almost been solved, right? That's kind of why you see things plateauing off a little bit as well. What's really interesting is how you make those language models do valuable things for enterprises, or how do you solve real problems? How do you deliver ROI? There's a lot of anticipation around, yeah, we need to show that there is a return on investment for all of these huge investments that have gone into AI.
Douwe Kiela [11:56] So the way to do that is to really build systems that solve the problems. And the model, the language model, is really a very small part of that much bigger system.
Controlling AI hallucinations at the model level
Matt Turck [12:26] As a great transition into this, maybe last question on sort of the general frontier model, large model kind of part of the discussion. And obviously a key objective of RAG is to reduce hallucination and all those good things at the model level. So if you take the Anthropic and the OpenAI models, are you seeing progress around hallucination control specifically? And if so, what are the drivers?
Douwe Kiela [12:57] So on the model specifically, I think models still hallucinate, and I don't even necessarily think that that's a bad thing. So first of all, hallucination is very ill-defined, right? I think a lot of people are kind of conflating it with being wrong. But I think hallucination is a very specific type of being wrong where you make up information that is not grounded in what you consider to be the ground truth. So I think for a general-purpose language model, if you deploy this in a marketing department or creative writing, then hallucination is a feature.
Douwe Kiela [13:34] It's not a bug, right? You want it to generate beautiful prose. Maybe you don't care that much about the factuality of things. So it's really a problem. The underlying problem is that we want these language models to be good at everything. So they have to be generalists. And so they also have to be useful for creative writing. And I think that's wrong. Where we're headed is that we will have more specialized language models. So we have our grounded language model that has specifically been trained to be grounded, and that model hallucinates much less.
RAG: definition and origin story
Douwe Kiela [13:51] So it's really much more strongly coupled to the context. And so it's not great for creative writing, but it's very good at RAG problems.
Matt Turck [14:21] So let's get into RAG. First of all, it'd be awesome if you could quickly define what that is for anyone that's still learning about the space. And then we'll get into all sorts of technical details. And then I'd love to go into the origin story of it. So you're the lead author on the RAG paper at FAIR that came out in 2020. I'd love to hear that story. So, definition and then the birth of RAG.
Douwe Kiela [14:45] Yeah, sure. So RAG is actually a very simple idea. It's answering a very basic question, which is: how do I get a generative model, so a language model, to work on top of data that it was not trained on? So RAG is really the way that everybody has gen AI work on their data. And the advantage of that is that you can always stay up to date. You don't have to constantly retrain the model when new information arises.
Douwe Kiela [15:18] And so because it's so easy and it's such an intuitive way to do this, it has really become the dominant paradigm. And so the story of the history of RAG, and this was just really a great team collaboration with lots of fantastic people involved. I originally became interested in this problem because I was interested in grounding, and my PhD thesis had been about grounding in different modalities. So vision or audio or even smell and taste, you can ground in all kinds of different perceptual modalities.
Douwe Kiela [15:50] So then I arrived at FAIR and I was thinking about multimodality and grounding. And then I was like, why can't we ground in other text? Why can't we ground text in other text? So we then need some source of truth. And so Wikipedia is the obvious answer, despite all its flaws. So we basically said, okay, let's say everything in Wikipedia is true. So now, when we generate answers, they can be grounded in what we consider to be true without actually having to be trained on that.
Douwe Kiela [16:22] If you update something in Wikipedia, that should be reflected in the answer without needing to retrain. So we got very lucky with that project because FAISS had just happened, which was basically the first vector database, Facebook AI Similarity Search. It's a great open-source library that is still getting used a lot these days. And so we had a vector database, basically the very first vector database, and we had a generative model, and we put them together, and then, yeah, things worked.
Douwe Kiela [16:34] So that was great.
Matt Turck [16:47] And you alluded to this, but how does it even work in a research organization like FAIR? How does one decide to focus on this topic or that topic and get approval? And how does that actually work?
Douwe Kiela [17:06] No, there are no approvals, but now there are approvals. At the time, this was the beautiful era of AI. I think I'm incredibly lucky to have been a part of that. There were no rules. I arrived as a postdoc, and I think I had six interns in my first year, and I could work on whatever I wanted. I was looking at Wittgensteinian language games and all kinds of weird stuff, emergent communication, multi-agent systems, and things that a lot of the things that I was looking at at the time, they're starting to sort of become cool now.
Douwe Kiela [17:47] But at the time, a lot of people thought that I was kind of crazy for even looking at these things. And yeah, so I was very blessed that they let me look at that stuff. And so RAG is really an example of that, right? It's just thinking of kind of frontier ideas, and then it turns out that they actually work. And so when the RAG paper came out, I don't think a lot of people were super excited about it. It happened much later.
Douwe Kiela [18:00] You can even see this in the citation profile on Google Scholar. It's like, oh, the RAG paper is nice. And then it's like, holy shit, RAG works.
Matt Turck [18:08] So there was no particular thinking behind the name or any of the things, because it's become such an industry term.
Douwe Kiela [18:32] I think Patrick Lewis, the first author, also often jokes on podcasts and things like, we could have come up with a much better name than this. But I guess it was a pretty decent name, otherwise it wouldn't have stuck. And so there was a lot of other interesting work happening at the time. So Google had a great paper on REALM. So that was more like a BERT-style masked language model. So it wasn't really generative, but they had a lot of the same ideas in there, and that also worked really well.
Why the Transformers paper initially felt underwhelming
Douwe Kiela [18:47] But why RAG became the way you name these things is because it's generative, right? So we were the first ones to have a generative model there. Yeah.
Matt Turck [19:11] And it's funny how these things happen, right? There tends to be something in the water at any point in time when different groups of very smart people in different organizations do think similarly. I think I heard you somewhere talk about how, when Transformers came out, from your perspective, that was actually sort of underwhelming. Can you go into that quickly?
Douwe Kiela [19:35] Yeah. I mean, it's just fun how history is sort of rewritten over time, I guess, by PR departments of large enterprises. But when the Transformers paper came out, at FAIR, we were working on very similar ideas, obviously, I guess, a bit more in the convolutional space because of Yann LeCun's background. And I guess that's what a lot of people were also exploring at the time. But the idea of the Transformers paper is really just, can I cut the recurrence of the RNN?
Douwe Kiela [20:05] At the time, we had RNNs, right? Can we cut the recurrence? Because then we can do parallel processing much more efficiently on a GPU. And so that turned out to work pretty well. You had to do a couple of tricks then because you lose your ability to understand positions. So you need positional embeddings, and they came up with some really cool ways to do that. And you ideally want to have the attention mechanisms sort of have multiple tries, so that became multi-head attention.
The core architecture of RAG
Douwe Kiela [20:41] And that is really just what the transformer architecture is. So I would say, and maybe I'm biased because one of my best friends is on the original attention paper, but that was the real breakthrough. It's just like figuring out that you have this attention mechanism that actually allows you to do a much better job at sort of representation learning and, as a result, kind of generating correct text sequences autoregressively.
Matt Turck [20:57] 4.0 version, and agentic RAG and what Contextual AI does. But maybe as an introduction/deep dive, how does that work? So there's a retriever, there's a generator, all those good things. What's the core architecture?
Douwe Kiela [21:32] Yeah, the core architecture is very simple. You have a language model, so that's the G, and then you want to give that context. And the way you do that is by augmenting it, the A, using retrieval, the R. So that's RAG. And so how you do the retrieval, that has been changing constantly over time. So in the initial paper, we used a vector database, or FAISS. The words "vector database" didn't exist at the time, so FAISS was the first vector database.
Douwe Kiela [22:10] And I think people over time have started figuring out that that has all kinds of limitations. So how a vector database works is you just have embeddings. So you encode pieces of information or chunks of documents, you encode them as a vector, and then you do basic dot-product similarity search. But that has issues where you're just looking for chunks that are similar to the question, but you don't necessarily want to find chunks that are similar to the question. You want to find chunks that are relevant to answering the question.
Douwe Kiela [22:47] So you need to do different things with the representations, with the embeddings. So a lot of modern RAG deployments are very different from the original ideas in the paper. So you still have a vector database, but the way you encode things is very different. Then you usually also have a sparse component. So it's not just dense search, it's not just vector search. You're also doing even older traditional keyword-based search, so BM25 or TF-IDF-style algorithms.
Matt Turck [22:49] Or Elasticsearch, or that kind of stuff.
Douwe Kiela [23:12] Yeah. So that's what Elasticsearch is very good at, right? And so then you have this hybrid retrieval system, and you can cast a relatively wide net with that. So it's very cheap to do that, but then it's not going to be great. So on top of that, you need to have a re-ranker that actually filters out most of the stuff that you probably shouldn't have retrieved in the first place. So you get this kind of cascade of retrievals where you narrow down the search, and the model that does the decision-making gets smarter and smarter the further in the cascade you go.
Douwe Kiela [23:51] So the final re-ranker step, that can be a pretty beefy model where you can even, in our case, give it instructions, which is awesome, right? So you can really tell it, like, I believe this source much more than this source, and I have a strong preference for recency. And if it's a PDF, then I believe it much more than if it's our internal Slack or something, right? So that type of ranking really factors into the overall retrieval results. And then the final step is that goes to a language model.
Douwe Kiela [24:14] And hopefully that language model doesn't hallucinate, but you have no guarantees. So that's kind of how the original naive RAG evolved into sort of advanced RAG, where you have this pretty sophisticated pipeline and it's really a system, right? These are all different models. And so one of the—
Matt Turck [24:23] To unpack some of this at a very practical level, then what goes back to the model is in the form of prompts. It pushes information into the context window.
Douwe Kiela [24:44] Yeah, exactly. So your retrieval results, the context that gets given to the language model as a part of the prompt, it's like, this is the question. And then you go off and ask the question to your retrieval system, you get the results, you put them into the prompt, and then you ask the language model to do its magic.
Matt Turck [24:58] And for the hybrid search, how do you, or how does the system decide when to do a term search versus a vector search? Is there presumably intelligence there in allocating certain types of query to certain—
Douwe Kiela [25:20] In most systems, there is no intelligence there. So in a lot of advanced RAG, that's sort of a hyperparameter that you just tune. So in our case, that is learned. And so it is something that you can learn, but we take a very different approach. So we kind of started from this observation that RAG is really not about models, it's about the system of models. And so the real question is, how do you get all these models to work together in the right way?
Douwe Kiela [25:55] So in our case, each of those components, right? So the language model has been trained to be grounded. So it's state-of-the-art at grounded generation. Then we have a re-ranker that has been specifically trained to work together well with that grounded language model. Then we have our retrieval step before that, right? And then all of these components are jointly optimized, so they're trained on the same data distribution, and that means that they have been trained to work well together. And you can really see the difference in benchmarks in terms of performance when you do that.
RAG vs. fine-tuning vs. long context windows
Matt Turck [26:32] Maybe as one final question before we transition to that part, you mentioned that RAG has become one of the dominant architectures, which is certainly, from my perspective as an industry observer—I mean, investor, but industry observer—that certainly feels like it's the case. A few months or a year or so ago, there was some level of debate. And if I had to summarize it, some people were saying RAG is great, but there's actually a lot you can do with fine-tuning. And then perhaps more recently, there was a school of thought that seemed to be saying RAG is actually going to go away because the context window of models has extended so much that you just can put the entire context into the window.
Matt Turck [26:50] Where are we on that debate today?
Douwe Kiela [27:04] Yeah, I think it's so fascinating. It's very related to the U.S. versus China observation, actually. It's like, somehow, it's journalists and venture capitalists, I guess, who really like these dichotomies, right?
Matt Turck [27:16] Those two categories increasingly tend to overlap as we create more and more content that the world really wants to see, clearly. No, it's amazing content.
Douwe Kiela [27:44] But yeah, so to answer your question, these are not dichotomies, right? So it's not RAG or fine-tuning, and it's not RAG or long context. It's all three of those things, ideally. So if you have a RAG system, you can actually fine-tune it, and that will probably be better. You don't want to just fine-tune the language model. You want to fine-tune the entire system, so also the re-ranker, also the embeddings, ideally as much as possible. So that basically just says, if you know what you're specializing for and you have a more specialized problem, then using machine learning, you can always do better on that problem.
Douwe Kiela [28:22] So the question is, do you want to invest the resources required for this? So you will need compute to do that, and you will need to spend time to actually make it work. But if you make that very easy for people, like we do, then you can have an out-of-the-box amazing RAG system. And then you can specialize that very easily to make it even better on the use case you care about. So then you get the best of both worlds. And so one common misconception about fine-tuning is a lot of people think that you can inject new knowledge into a model using fine-tuning, and that is not true.
Douwe Kiela [29:00] So you can make existing knowledge sort of come out more, or you can make it adopt a specific style of communication or something like that. But telling it, like, actually, this thing I taught you, or this piece of information that you were pre-trained with is no longer true. Now, Trump is the president instead of when you were trained, it was Biden. Those types of things you really cannot get into a model with fine-tuning easily. So you need RAG anyway, and you want to do fine-tuning anyway.
Douwe Kiela [29:31] So that's not a dichotomy. And then the third one is about long context, where my example is always, like, a basic question, like, who is the headmaster in Harry Potter? So it's Dumbledore, and you didn't have to read all seven books to figure that answer out, right? You could have just searched for headmaster, maybe, or read chapter one, and then you would get the answer. So long-context models are inherently incredibly wasteful. You're paying for all this compute, and that's maybe why some of the companies that are trying to really sell long-context-window models, they will make more money from that, right?
Douwe Kiela [30:05] Because you're spending more on compute for reading Harry Potter for every single simple question. So you want to have a long enough context window so that you can fit in the relevant pieces of information so that the language model can do its job. So if you have RAG, you can do this over millions or trillions of documents and everything will still work. And then you just pick the pieces of information that are relevant and give them to the language model as context.
Douwe Kiela [30:40] So again, there is no dichotomy here. You probably want both. And even for different types of problems, you want to have different solutions. So if my problem is, summarize this document, then I will probably want to put it in a long context as much as I can. If my problem is, like, what is the voltage of pin seven on chip X for this particular manufacturer, and how does it compare to this other one? That's a RAG question, right? You should just retrieve that information, do the comparison, reason over it.
RAG 2.0: Thinking in systems and not models
Douwe Kiela [30:55] So, yeah, they are just kind of different solutions for the same underlying problem, which is you have more information than you can feasibly sort of spend compute on.
Matt Turck [31:21] So, as you alluded to, the fundamental idea is to think in systems and not in models. So, the way I understand it, the general idea is to have all components of a RAG architecture, instead of a Frankenstein kind of assembly, having all those components deeply integrated and learn together. Can you describe, I guess, in better terms than mine?
Data extraction and data curation for RAG
Douwe Kiela [31:50] Oh, that was a great description, actually. I mean, so that's really the basic idea, right? So it's really like starting from the idea that it's a system. And so, by the way, that also includes extraction, right? So there are lots of very interesting problems in terms of document understanding where existing solutions really fall short. And so if you want to have an enterprise-grade RAG system, you are only as good as the data that goes into that RAG system. So if you can't extract the data in the right way, so if you have a sort of table structure and it has nested information, if you can't get that out in the right way, then your RAG system is going to fail completely.
Douwe Kiela [32:05] So that is a part of—
Matt Turck [32:06] Garbage in, garbage out.
Douwe Kiela [32:16] Exactly, yeah. So it's beautiful data, beautifully formatted PDFs, very easy to read for humans. And then AI can't do anything with that information because it can't get it out.
Matt Turck [32:30] Is that a human challenge or is that a technical challenge, or both? And by human, I mean internally, the enterprise deciding what information goes into the vector database.
Douwe Kiela [32:55] Yeah, so it shouldn't be a human problem. Sometimes it is now for some companies, but it shouldn't be, right? So you should just have an AI system that can just read the document. That's not that much to ask of an AI system, right? But understanding PDFs is a great example, right? How humans understand PDFs is by looking at it visually, but how a machine understands it is by looking at the actual encoding, right? So the raw information that makes up that PDF.
Douwe Kiela [33:25] And so what we're doing is we're actually looking at the PDF the same way that a human does. So we're looking at the entire layout. We have a layout segmentation model that says, hey, this is a chart, this is a table, this is a piece of text. And if it's a table, then it's extracted differently using a different table extraction model. And if it's a graph, then it's extracted differently because that's a different modality for the data. So all of that information is then put together in your extraction output, and that is what ends up in your retrieval database for your RAG system.
Matt Turck [34:01] But the reason, just to double-click on this, who decides that the PDF goes into the vector database in the first place? Is that a complicated sort of enterprise discussion? And I guess the parallel to that is, is there a world where at some point the RAG system can just go fetch the information wherever it is versus having one database that's considered the RAG database?
Douwe Kiela [34:22] Yeah. So that world is already here, I think. So there are different ways to do that. One is to make sure that you have access to all of these different systems, and then you synchronize that with your vector database. Or the other is to just have your language model reason and use tools and actually call into other APIs. So I don't have to index all of Slack if I can call Slack's search API and get the information that I need.
Douwe Kiela [34:54] So that has benefits because then I also don't have to worry about entitlements or role-based access control or things like that. So there are different strategies you can follow there. But I think you're right, like, for enterprises, that often really is just a human decision. It's like, what data goes into these platforms and how do I control that data? That's a really important problem. So just giving the system access to all data in a company is often not really the right way to do it.
Douwe Kiela [35:23] But ideally, you should be able to put any data that you want into it and then expect it to work. And that actually is often not true, right? So real-world data is very noisy. You can build a very awesome demo on a couple of PDFs and things will probably work, but then you have to scale it up to a million PDFs and then everything breaks down. And the reason for that is that a lot of these kind of advanced RAG systems still actually don't have this re-ranker working well enough to resolve a lot of these data conflicts.
Douwe Kiela [35:54] And because the parts are completely disjoint, right? So the retriever and the re-ranker and the language model, they're completely unaware of each other. And so our approach, where everything is much more tightly integrated, means that you can be much more robust to scaling it up to much more and much noisier data.
Contextual Language Models (CLMs)
Matt Turck [36:01] So that's the extraction piece. We alluded to the re-ranker piece. CLM architecture?
Douwe Kiela [36:38] Yeah, so for us, that's really—we need to have a language model that is grounded much more strongly than a general-purpose language model because we have a very low tolerance for hallucination. And that's because we are an enterprise company and because we're focused on RAG. And with RAG, you want to give true answers. That's the whole idea of RAG, right? So, yeah, we found that standard language models are just not good enough for what we need. And that's why we've trained language models to be much more strongly grounded and to be much more contextual and really look at what's in the context.
Douwe Kiela [37:02] And anything that is not in the context, it will just say, "I don't know," which is one of the superpowers of our system: the ability to say, "I don't know." That's really what you want if you're in a high-stakes setting.
Matt Turck [37:08] And are the CLMs based on open source, or are those models that you started from scratch?
Douwe Kiela [37:28] Yeah, so we initialize from open-source components. The main one we use—we have some flexibility there; it depends on the customer—but the main one we use is Llama. And so we initialize with Llama and then do a lot of training on top of that to make sure that Llama is actually grounded, because the original Llama that we start off with is not all that grounded.
Matt Turck [37:34] Another great contribution by Meta and Facebook to the AI world.
Douwe Kiela [37:37] Yeah, Llama and PyTorch. Yeah, so much.
Matt Turck [37:53] Yeah, it's sort of amazing, as an aside, the impact of Facebook originally and Meta now on technology in general is something that not everybody realizes, right? Like, if you had React and—
Finetuning and alignment techniques: GRIT, KTO, LENS
Douwe Kiela [38:04] Yeah, they don't get enough credit, honestly. The world would be a very different place without PyTorch and React and Llama and things like that. I think they're doing amazing things.
Matt Turck [38:11] Let's talk about some of your fine-tuning and alignment techniques: GRIT, KTO, LENS. What are those?
Douwe Kiela [38:42] Yeah, so KTO is a different way to do DPO, so direct preference optimization. So RLHF, when that came out, reinforcement learning from human feedback, everybody was like, oh, we have these preferences. So essentially, we have maybe two different possible answers that the language model might give. And now we want to train it to say, actually, this one is more preferred by humans than this one. So this was originally kind of the secret sauce behind ChatGPT. That's why ChatGPT suddenly really started working, because it captured people's preferences.
Douwe Kiela [39:15] But the problem with RLHF is that you need to have this very heavy reward model, and that needs to be trained up on these pairwise preferences, or you need to have multiple of these generations, which is really problematic. So DPO then said, okay, actually, we can do some smart math, and then we don't need the reward model anymore, which was great. You can directly optimize on the preferences. But then when we were thinking about this, we were like, that's still not ideal, because in the real world, you don't want to collect a thumbs-up for every thumbs-down.
Douwe Kiela [39:50] Like, when you generate an example and somebody tells you, actually, that wasn't a good example, then you want to be able to learn from that without being told what the right example was, or vice versa. When I get, like, oh, that was great, should I then go and generate a couple of bad versions of this in order to train on it? That just doesn't make sense. So with KTO, we were like, can we directly optimize on the feedback without there being preference pairs?
Douwe Kiela [40:29] And that turned out to work really well. We've done some work since on anchored preference optimization as well, so APO, and that really is the best direct preference optimization technique that we've seen. And so it has all kinds of benefits around not having to rely on preference data. You can just incorporate direct feedback data. And it's very, very sample efficient. So you can train on this when you only have, like, 100 examples. You can really make a meaningful difference. And that's great because data annotation is very, very expensive and very cumbersome.
Agentic RAG
Douwe Kiela [40:40] So the more you can make the system learn just on its own, the better it is.
Matt Turck [41:16] Another very interesting part of what you all do at Contextual AI seems to be around agents, and that gets us into agentic RAG. And that was very eye-opening for me when I was prepping for this. But when I first came across Contextual AI a few months ago, my mental model is that one can think of a retriever as a tool as part of an agentic workflow, RAG being a specialized sort of agent use case. But is that the right way to think about it?
General vs. specialized RAG agents
Douwe Kiela [41:37] I think so. So what we have, our product is a platform for building agents, but you need these agents to work on your data. And the way you do that is through RAG, right? So we're a platform for RAG agents, and you can specialize these RAG agents for different use cases using the platform very easily.
Matt Turck [41:45] You have a general RAG agent, and then you have specialized RAG agents. Is that right? So why is that, and who does what?
Douwe Kiela [42:04] Yeah, so just the starting point is really you can create your own agent in minutes. It's very easy to do that. You don't have to do much. And out of the box, that will already be much better than what you could have built yourself, probably with a kind of open-source RAG framework. And then you can specialize that using machine learning. So you can use our Tune API to really make all the components optimized for the specific problem that you're trying to solve.
Douwe Kiela [42:29] So that's how you get to specialized RAG agents, where you can really hit the production bar and actually deploy this in a setting where you have much more control over what is right and wrong. You don't just have prompting; you can actually train that entire system to be good at what you need it to be good at.
Matt Turck [42:53] And you train them in the way one would train a regular AI agent? Bearing in mind that nobody really knows what an actual AI agent is today, or maybe you do, but I certainly don't. Where I'm going with this is the concept of planning, of failure mode, retrying, kind of sharing your work transparently, all those things. Is that part of what you do?
Douwe Kiela [42:54] Yeah, that's part of it.
Matt Turck [42:54] Yeah.
Douwe Kiela [43:21] I mean, we're also working on that like everybody else, right? So I wouldn't say that journey is finished in any way yet, but yeah, that is really where everything is headed. So for us, a RAG agent is really about: you get a question, and then you formulate a plan about where am I going to get this information from, right? It's like, maybe I need some unstructured data, but actually I also need to query this database. So we also support structured data, which is very important.
Douwe Kiela [43:50] And then maybe I need to call this API, right? So you formulate a plan, or what we call our Mixture of Retrievers figures that out. So that's kind of an intelligent retrieval strategy. And then that goes to our reranker, and then that goes to our Grounded Language Model. So where this is all headed is really in this test-time compute paradigm, where, like you said, retrieval is just one of the many tools. And you can decide for yourself, like, actually, I retrieved this thing, but it's not exactly what I'm looking for.
Douwe Kiela [44:08] Let me try this. So yeah, that's all very much happening. And that's how agents are going to be valuable on your data, which is really where we need them to work.
Matt Turck [44:19] And if you crack that, is there a world where Contextual becomes an agent platform, full stop? Not just for RAG, not just for information, but taking any action in the enterprise?
Douwe Kiela [44:34] I think most of the interesting problems are actually RAG problems. That's why it's such a dominant paradigm, right? Like, an agent that doesn't work on your data is not very useful. So if you need it to work on your data, you probably need a RAG agent to begin with.
Synthetic data in AI
Matt Turck [44:49] One more topic I thought was interesting while prepping for this was synthetic data. It sounds like your thinking has evolved into how helpful it is, especially in the context of RAG. If you could go into that.
Douwe Kiela [45:17] Yeah, I think synthetic data—so O1, and I guess DeepSeek, showed that synthetic data is actually pretty valuable, right? And we have known about that for a while because that is a great way to actually train up a RAG system end to end. So you can use synthetic data to make all of your components better, and you can also make sure that the components work well together through some really clever tricks with synthetic data. So a lot of our joint optimization is just leveraging that observation.
Deploying AI in the enterprise
Douwe Kiela [45:51] And that's also how we can be very good in terms of specialization. So ideally, you want to be able to specialize without needing any data annotation. So you don't need to have any examples of what right and wrong looks like. You just tell me what you want, and then I will figure out what right and wrong looks like and train on that. And so that is really unlocking lots of very cool opportunities, I think, for these specialized RAG agents.
Matt Turck [46:22] To close the conversation, I'd love to go into a little bit of your experience selling AI in the real world, in the enterprise. There's a lot of cool stuff on Twitter, a lot of cool demos, but it feels like you're exactly at the spot where sort of the rubber meets the road in terms of actually deploying AI for real-world cases in the enterprise. What have you seen in terms of where we are in terms of adoption and opportunities and obstacles? And where do you think we are in the cycle?
Douwe Kiela [46:57] Yeah, I think we're still pretty early. There has been this sort of first wave where everybody wanted to do everything themselves. And so we have had lots of sort of build-versus-buy hurdles to overcome there. I think people are starting to realize much more that you cannot build a sort of RAG platform like ours yourself and then keep maintaining that forever and sort of keep up with all the latest innovations and trends that happen continuously. You see how quickly things change in AI, right?
Douwe Kiela [47:29] So I think the market is sort of waking up to that observation much more. And in terms of maturity, I think a lot of folks initially were thinking about it as sort of like, we need to get something in production, right? So it was kind of aiming too low. It's like, oh, we do some internal enterprise search, and now we can ask who our 401(k) provider is. That's great. You could probably already do that before GenAI. But that's not where the ROI comes from.
Douwe Kiela [48:06] So that's why we're so focused on these really high-value knowledge worker use cases where you have knowledge professionals. If you can make them even 10% better at their job, you can save companies millions or hundreds of millions of dollars by doing that successfully. But doing that is a very different game because real-world data is very noisy. You need to specialize if you really want to solve these problems at a sort of level of accuracy where it's actually useful and delivering real value.
How tolerant are enterprises to AI hallucinations?
Matt Turck [48:19] What's the level of tolerance you've observed for hallucination or inaccuracy for this non-deterministic system?
Douwe Kiela [48:32] Yeah, so that has been evolving. I think initially we had people ask us, like, when are we getting to 100% accuracy? And I had to give them the bad news that probably never. But just like with humans, right? Like in the financial sector, there's a reason we have regulators, and there's a reason we have very stringent processes around what people are allowed to do, and how we kind of control that, and how we double-check that there aren't any mistakes that crash the market.
Douwe Kiela [49:13] So, in a very similar way, we need to look at the inaccuracies of AI systems and then work on ways to mitigate the risks. And so, that is something I'm really seeing much more of now, especially in financial services. In healthcare, where a small mistake can have huge consequences, you just want to think very carefully about the inaccuracies. The accuracy is sort of table stakes, but the inaccuracy, that's really where you need to think about ways to handle that.
The future of Contextual AI
Douwe Kiela [49:36] So we do things like providing very fine-grained audit trails for regulators, making sure that we verify each of the claims that the system makes and that we flag things if the system is not sure about its answer, things like that.
Matt Turck [49:39] So what can we expect from Contextual AI in the next year?
Douwe Kiela [50:01] Yeah, more progress on this mission to change the way the world works. That's ultimately what we're trying to do. So literally how people do their jobs, right? So yeah, the big things we're focused on are the intersection of structured and unstructured data. I think that's really exciting. There's so much cool stuff to do there. Multimodality is obviously a big theme, and test-time reasoning, but then really focused on retrieval because that's really the only way you get these agents to work on your data and your problems.
Matt Turck [50:21] Douwe, thank you so much. Terrific. I really appreciate you sharing all of this with us, and excited for the next year.
Douwe Kiela [50:22] Thanks so much for having me.
Matt Turck [50:44] Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already, or leaving a positive review or comment on whichever platform you're watching or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.