OpenAI's Yann Dubois: Why AI Progress Suddenly Feels Real

The MAD Podcast with Matt Turck · with Yann Dubois, Co-Lead, Post-training Frontiers at OpenAI

Yann Dubois is the Co-Lead, Post-training Frontiers at OpenAI. We cover why coding tools feel suddenly useful only after crossing a reliability threshold, how reinforcement learning moves reasoning from verifiable competition problems to real-world utility, and why startups’ biggest opportunity is the last mile of permissions, connectors, and vertical workflows.

Watch on YouTube

Chapters

  1. 1:30 — Why recent AI progress feels like a step function
  2. 4:13 — Model reliability & the rollercoaster of shipping 5.5
  3. 7:33 — How OpenAI structures vertical and horizontal teams
  4. 9:49 — Improving model efficiency and test-time compute
  5. 12:32 — Yann Dubois' journey from Switzerland to OpenAI
  6. 15:37 — Reasoning in 2026: Real-world utility vs verifiable rewards
  7. 18:34 — GPT-5.5 Thinking vs Pro: Scaling test-time compute
  8. 20:09 — How reasoning models become more efficient
  9. 23:23 — Pre-training scaling and overcoming the data wall
  10. 27:03 — Multimodal data, synthetic data, and embodied AI
  11. 31:05 — Demystifying mid-training and post-training
  12. 37:21 — Does RL create new capabilities in AI?
  13. 38:53 — The challenges and frontier of scaling RL
  14. 43:09 — Is building AI models a craft or a strict science?
  15. 48:21 — How AI models generalize across different domains
  16. 54:18 — How reinforcement learning cures AI hallucinations
  17. 56:04 — Negative generalization and conflicting instructions
  18. 58:05 — Can RL scale to law, medicine, and the broader economy?
  19. 1:00:19 — The evaluation bottleneck and Model as a Judge
  20. 1:04:21 — Continuous AI progress & continual learning
  21. 1:08:49 — Will foundation models eat the agent harness?
  22. 1:11:23 — Why startups should focus on the last mile of AI

Transcript

Why recent AI progress feels like a step function

Matt Turck [1:23] Hey, Yann, welcome.

Yann Dubois [1:30] Hi, Matt. Thanks for having me.

Matt Turck [1:50] GPT-5, Claude, Mythos Preview. So it feels like we have unlocked yet another step function in progress, particularly in cybersecurity, agentic coding. What's the best way to think about this from your perspective? Are things accelerating? What is happening?

Yann Dubois [2:27] Yeah, the last few months have been pretty wild. Internally, we also really feel it. And I think anyone who's coding basically is really feeling it right now. I think that's really because of three reasons. The first one is, even though in my mind the progress is actually pretty continuous, you need to reach this level of reliability to really make any of these AI tools very useful. And I think we just crossed that, probably December last year, at least at OpenAI. That's where I thought we really crossed that threshold, where now we can trust these models to do a lot of the work that we are doing.

Yann Dubois [3:08] So it feels like a step function, even though I think actually in terms of capability, it's pretty continuous. So that's the first thing. The second reason is once you start having models that are really good, you accelerate yourself, especially in terms of coding, given that we all code internally. You accelerate yourself both for having these models train the other models, but also build the tooling that we need as researchers to do our job. And all this acceleration, I think, means that we saw these last few months going faster and faster.

Yann Dubois [3:46] The third thing that I think we are feeling is all of last year, we really built these reasoning models and we really saw pushing a lot on reinforcement learning. And initially, when we had, like, o1, o1-preview, even o3, these models were still optimized for what we call verifiable rewards, things where we actually have access to ground truth and it's easy to test whether you're correct or not. That is, for example, the case in math questions or coding competitions. And what I think we are realizing now is that we were able to take many of the tools that we built for these verifiable reward cases.

Model reliability & the rollercoaster of shipping 5.5

Yann Dubois [4:14] And we were able to use them more generally for reinforcement learning on real use cases. And I think that's really why we're feeling that right now in just real-world coding rather than competition. So we moved from competitions to usefulness to users, and that's what we are feeling right now.

Matt Turck [4:27] Okay. Fascinating. So we're going to unpack a lot of this, particularly on the RL side. First thing that you mentioned: reliability. Is that in engineering? Is that models? What makes a model reliable in the way you meant it?

Yann Dubois [4:46] It's a little bit of everything, but in general, given that these are agentic models, the longer—if you just think about it as every two minutes, there's a certain probability that they're wrong. The longer that they run, the higher the probability that the final answer is going to be wrong. So it's just something inherent in agentic models. And what we've been pushing a lot on is making sure that we decrease this probability of being wrong every two minutes.

Yann Dubois [5:12] So purely from a model point of view, of course, there's a lot of reliability that is also happening on the applied side. And the team at OpenAI has been doing an amazing job on that. But I'm even talking only about reliability of our models and basically making sure that we decrease the probability of being wrong. Great.

Matt Turck [5:33] GPT-5.5, formerly known as SpUD, was, as mentioned, a big deal—is a big deal. And I'm just curious, from the inside, what are you guys the most proud of? What did you find the most challenging? Give us some color on how you all felt releasing this.

Yann Dubois [5:52] GPT-5.5, to be honest, is one of these models where everyone in the company was extremely involved in building. And I think that we really feel it now. GPT-5.5, it seems like all the stars were aligned. That doesn't always happen. And that was just a great model for this. So we did feel it. It's kind of funny because in general, with every model that is looking really good early on, we have a model, we all get really excited about it, and then tons of doubts start coming up because it's like, oh, everyone is hyping this thing internally, but actually it's bad at all these other things.

Yann Dubois [6:38] And then there's another wave where people start underhyping it, and it kind of goes through waves. And it depends when we actually ship it, how people feel about it internally. But that's true of most models that we have. GPT-5.5 was not that different in this case, but it definitely maybe had a higher amplitude of the wave. So people were very excited, then not as excited, and we shipped it, and people were happy externally.

Matt Turck [6:56] How long does that process take, including the waves of going up and down and excitement? I guess it depends on the release and the importance of each release. But is that a few weeks? Is it a few months?

Yann Dubois [7:25] It really depends. GPT-5.5, but it can—it kind of depends which part of the pipeline is training parts of the model. So we really have different sub-teams, including pre-training, and you have the mid-training stage, and you have some post-training. And usually, the closer you get to products, post-training being the last one, the faster the iteration cycle is. And if you're more upstream, the slower the iteration cycle is.

Matt Turck [7:29] Okay.

How OpenAI structures vertical and horizontal teams

Yann Dubois [7:33] So it could go from, let's say, from months to days, basically.

Matt Turck [7:49] GPT-5 was particularly good on agentic coding, computer use, knowledge work, and early scientific research. How does that work internally? Do different people focus on those different parts? How do you get to that result?

Yann Dubois [8:13] Yeah, we definitely have different teams that are working on specific use cases and are pushing on these use cases. My team specifically is actually the one that is taking all these vertical improvements and trying to put them together in the final model. You could see it as a team that is doing both the smoothing function. So you have all these improvements, but you need to make sure that the model doesn't feel too spiky, doesn't feel differently on different verticals. And also, you need to have some teams that are working—and that's basically what my team is doing—on all the horizontal improvements.

Matt Turck [8:20] Okay.

Yann Dubois [8:47] Improvements. So there are many things like instruction following, function calling, or thinking about how much should a model think for different problems. Those are very horizontal, and that impacts all these use cases. So we have both these more vertical teams and these more horizontal ones, and both are very important to improve on the model. And the good thing is that these things can kind of be improved orthogonally. So you might have multiple different teams that are working on certain verticals.

Yann Dubois [9:17] And maybe for one model, there's only half of these teams that made integrations, basically, in the last run and improved the model on these capabilities. And maybe for the next model, it'll be the other half. So that's kind of, at a high level, how it works. One thing which I will say, because you asked also about what are the things that we are really proud about for this model, I would say two things. Number one is the efficiency of the model.

Yann Dubois [9:47] We really, really improved the efficiency of the model, and most of the tasks can be basically performed, I would say, like, 2x faster now with this model. So that's great. And the other one that I already mentioned before is this alignment of the company and making sure that everyone is working towards the same goal. And that really takes the entire company working towards this North Star of building one good model in specific timelines. So very proud of how that happened.

Improving model efficiency and test-time compute

Matt Turck [10:02] Great. And then speaking of efficiency, how do you optimize for that? Are we talking about efficiency per token? Are we also talking about latency in serving the model? What part is AI research versus engineering?

Yann Dubois [10:29] So that's what I mean when I say it's the entire company: it really comes from everywhere. It has to come from inference optimizations. It has to come from the model being more efficient in its thinking time. So you have basically every token that you think for—the usual plot that you should be looking at is x-axis, the number of tokens that you think for, and y-axis, the performance. So these are these test-time scaling curves that we look at.

Yann Dubois [11:00] And research basically tries to move this curve to the left. So think less to be at the same level or more correct. And then inference also deals with this x-axis, but switches it from number of tokens to actual latency. And the final thing that people care about is latency on x-axis, performance on y-axis. And this is where everything comes together. So yeah, that's why I always say I'm really proud of the company for this one.

Matt Turck [11:14] Okay, great. Let's talk about you for a minute. So you are in the Post-Training Frontiers team. So that team you described as horizontal. So what does the team do in general?

Yann Dubois [11:28] Yeah, I would say there's three things that we do. So in a broad sense, we are under the post-training org, and my team is the Post-Training Frontiers one. So there are three things that my team does. Number one is we decide what goes into the final run. So as we talked about before, there's many verticals, and someone needs to decide what can go in, what cannot, and also provide the science experiments for people to iterate on something that's going to be representative of the final run.

Yann Dubois [12:03] So this is the first thing that my team does. The second thing that my team does is bringing everything together and actually doing the big run. As you might imagine, we train on a good amount of GPUs, so there's a lot of infra work that is needed, but also there's a lot of ML work that is needed by putting everything together and making sure things work well together. And then the third thing that my team does is horizontal improvements to the models.

Yann Dubois' journey from Switzerland to OpenAI

Yann Dubois [12:32] Basically, there are some things that these vertical teams will not usually look too much at. For example, the thinking time, as I said before. So, how much should a model think for on certain answers? Or instruction following, function calling, things like memory, and general improvements to the model that are really across the stack. So that's what the Pushing Frontiers team does, and I'm leading that team.

Matt Turck [12:35] Okay, great. And what was your journey into OpenAI?

Yann Dubois [12:59] Oh, it's a long story, but I'll try to keep it really short. Basically, I did my undergrad in biomedical engineering in Switzerland. I'm from Switzerland. And then I went on exchange in Canada, and I learned about Word2Vec. So, I don't know if you heard about this algorithm, but it basically takes words, which are something discrete, and puts them in a vector space. So, basically, a way to think about it is as a plane where words that are more similar to one another will be closer to one another.

Yann Dubois [13:39] So it brings these discrete words into some continuous space that is semantically meaningful. And I was absolutely blown away by that algorithm. And that's when I decided that I wanted to work on natural language processing and just understanding language. At that time, I was very wrong, but I thought that English NLP was basically solved or close to being solved. That was in 2017. So that was right when transformers started. And it was actually right before transformers. So I was very wrong, but I decided that I wanted to work on under-researched languages, and basically, I wanted to improve NLP on languages where we don't have that much data.

Yann Dubois [14:18] So I went to work for Grab in Singapore, and I was basically building the natural language processing pipeline for them, working with Khmer, Bahasa, Thai, Vietnamese, and all these different languages. And then I'm skipping a little bit. I did more academic-type work in different countries, and I ended up at Stanford, did my PhD there. And after this, had a small stint in startups and then went to OpenAI.

Matt Turck [14:30] Yes. And I remember seeing on your blog or your page a note for quant firms to not reach out to you because you were not interested in hedge fund work.

Yann Dubois [14:44] Yeah, but I always think it's very important for me to think about the positive impact that I'm having in the world, or at least that I'm trying to have. So that's why this thought is there. Yes.

Matt Turck [15:07] And as we were saying just before we started recording, people may have seen you in the GPT-5 video announcement, and you did this very funny demonstration of an app that was built on the fly to teach your partner how to speak French. So people should go check that out.

Yann Dubois [15:17] Exactly. That was a fun one. GPT-5 was not that reliable, so I was a little bit stressed that it wouldn't work, but it ended up working.

Matt Turck [15:24] So this was truly live and presumably very rehearsed, but truly live.

Reasoning in 2026: Real-world utility vs verifiable rewards

Yann Dubois [15:37] Actually, right before we did that, like the last rehearsal, it did not work. So I got slightly stressed about that, but yeah, seems like live ended up working well.

Matt Turck [16:11] Yeah, no pressure, but yeah, that landed perfectly. Okay, very cool. All right, so let's unpack some of the things we alluded to in the intro. So we started effectively talking about reasoning, and I'm curious what reasoning means in 2025, and also my experience as a user is that it's particularly good with messy data, which seems to imply that it needs to reason through ambiguity more. What has changed?

Yann Dubois [16:48] O1-preview were really breakthroughs in the research community about having models that can think. And the longer they thought for, the higher likelihood they would be of being correct. So that was really a breakthrough. But initially, and if you look at old blog posts, you would mostly see math evals and also maybe coding competitions, but things that are really easy to test whether you're correct or whether you're not. And that also gives you some suggestion about how we were training some of these models.

Yann Dubois [17:21] And how I see maybe all of last year, and especially the end of last year and the beginning of this year, is that we were able to take these algorithms that work with verified rewards, like things where we can say you're correct or you're not, to the messy real world and really optimize for the utility that we provide to users and making them more productive. So I think that's what really changed.

Matt Turck [17:30] Okay. So it's the post-training reinforcement learning part, largely?

Yann Dubois [17:54] Yeah, I would say that's—I mean, there's also another big part of it. Number one, basically the first thing is that, of course, when you develop a new method, the method is fragile and it's not that reliable, and it's hard to productionize. So this part also improved a lot. But then it's also really, basically, we had a tool that we could start optimizing for different things. And initially, when we were developing this tool, we were making a lot of simplifying assumptions in the real world, basically.

Yann Dubois [18:20] And now we are removing these simplifying assumptions and, at least in post-training, we are able to optimize really for user utility and make sure that these models are useful. And the tasks that we're looking at are useful. And that's why also now current evals look much more realistic. I mean, if you think about GDPval, or even if you look at SWE-bench Pro or SWE-bench, these look way more realistic than, let's say, some Codeforces or coding competitions that we were looking at with O1.

GPT-5.5 Thinking vs Pro: Scaling test-time compute

Matt Turck [18:43] GPT-5.5 Pro? Is that just more test-time compute, more tokens, and more time invested in solving a problem?

Yann Dubois [19:21] Yes, basically it's just a question of how much test-time compute we pour into the model, or we pour into this entire system that we're shipping. We've seen again and again: the longer the model thinks for, the better answers we will get. The problem is that these curves that we are talking about are definitely not linear. There's some plateauing effect, and they kind of look logarithmic in some sense, depending on which evals. So you can pour two times more compute and actually only get small performance gains.

Yann Dubois [19:56] I personally don't use Pro that much because I really don't like waiting. I'm pretty impatient, so I don't like waiting for that long. And I know that the probability of being correct definitely improves, but it doesn't improve enough for me to use it. But there are some people who use Pro and who really love it, especially for academic research. And I know a lot of mathematicians who are using it. And that's because they kind of just have this in the background that is running for maybe one hour, two hours, and they don't really need to iterate really quickly with the model.

How reasoning models become more efficient

Yann Dubois [20:10] And Pro is really good for that.

Matt Turck [20:33] I'd love to reconcile this with what you were mentioning about efficiency earlier per token. So is the idea that you would be able to think longer, but also be more efficient, therefore solve the task better? How do the time aspect and the efficiency aspect interact?

Yann Dubois [20:56] Yes. So if you go back to the plot that I was talking about, where on the x-axis we have latency and on the y-axis we have performance, we're basically moving this curve more and more to the left when we say that we improve efficiency. So we're becoming more efficient, or we spend less time to achieve the same performance. But what Pro does is that it extends this curve. So it says, "I'm going to think for much longer, but I will have a higher likelihood of being correct."

Yann Dubois [21:25] But every iteration of the Pro model also moves to the left. So it also becomes more and more efficient. The important part is there will always be tasks where you just want to maximize the probability of correctness and you don't really care about latency. For example, if I start a job before going to sleep, the model has like eight hours. It should just think for as long as it can.

Matt Turck [21:25] Yeah.

Yann Dubois [21:28] And this is what kind of probably gives you—

Matt Turck [21:42] In layman's terms, what does that mean practically? Or how does that work practically? If the model goes in the wrong direction, then it would interrupt itself earlier. Is that one of the axes?

Yann Dubois [21:47] So for the efficiency, okay, so there's two things. Are you asking for the efficiency? What does it mean?

Matt Turck [21:56] Yeah, for the efficiency. Yes, largely for the efficiency. I'm just curious how reasoning gets— Yes, that's a good question.

Yann Dubois [22:21] Let me give you maybe a metaphor from humans. If you have someone who's an expert in a certain domain and you compare them to some undergrad that is starting in that domain, the undergrad doing that task might take like one day, two days, and will have to think through a lot of the possibilities and investigate because it never did a certain problem. While someone who's an expert in that field will usually just know what direction to take, and it will not spend the time investigating ten different directions because it knows that there's one that is more likely to be correct.

Yann Dubois [23:03] So this is the type of efficiency that we're talking about. It's basically models where we optimize more on real-world problems. And as a result, it was trained to figure out with a higher likelihood which paths of reasoning are more likely to be correct. So this is the part on efficiency. There's also what you suggested: part of it is the model knowing when it's going down the wrong path. But this is also something that the model can be trained for with reinforcement learning.

Yann Dubois [23:22] It's knowing, okay, that seems like not a great path. Let me backtrack and let me go on and test something else. And if you train the model less, it might realize that it's on the wrong path much later.

Pre-training scaling and overcoming the data wall

Matt Turck [23:53] Okay. All right. So it seems like a lot of this goes back to reinforcement learning and post-training. So let's talk about how the different components of modern AI systems work. So let's talk about pre-training, mid-training, and then post-training, and spend more time on post-training since it's so important. Pre-training specifically: the big narrative of last year was that pre-training was hitting a wall and was not going to yield much progress. That seems to not be the case at all in 2026.

Matt Turck [24:07] Can you walk us through some ideas for what is happening in pre-training and why it's progressing now in a way that people hadn't predicted last year?

Yann Dubois [24:49] For pre-training, I can't talk in a lot of detail about what is happening internally, besides that the team has been doing a lot of good work and our models are really getting better and better. One thing that I do want to highlight when we're talking, for example, with efficiency: if you have larger models, the amount of thinking time, so the amount of tokens they will think for, will usually decrease. And the way that you can think about it is that, metaphorically, the model already thinks through its weights when it generates a certain token.

Yann Dubois [25:32] So you can decrease the number of tokens that it needs to generate for thinking by increasing the size of the model that you are training. So oftentimes, if you just increase the model size, if you basically pre-train larger models, you will get better efficiency. And the good thing with larger models is that they can be parallelized better at inference time. So even though you might think, okay, you actually generated fewer tokens, but by a larger model, so you actually might decrease the overall efficiency of the system.

Yann Dubois [26:08] This is not true, because the larger the model is, the more chances you have to actually optimize for inference on GPUs. So you will be able to make the overall system more efficient. So that's one thing I wanted to say with larger models, that they are actually giving you a lot of efficiency. Otherwise, in terms of pre-training, I think it's very interesting. I actually also thought maybe two years ago that pre-training was kind of hitting a wall. And when we see, for example, if we talk just about Anthropic, I mean, Opus seems like clearly just a much bigger model. When you look at the cost, the cost of the model, usually that's how you know.

Yann Dubois [26:53] By the way, if it's a bigger model, you just look at the cost per token. And clearly they're getting very good performance just by increasing the size of the model. So I think the field was very, at least part of the field was surprised about that. There were a lot of conversations about hitting data walls, and it seems like we did not quite hit it. So the larger the model is, the more data it needs to ingest to be trained. And it seems like different companies kind of found different ways to overcome the fact that we don't have that much data on the internet.

Multimodal data, synthetic data, and embodied AI

Matt Turck [27:12] Is the next frontier, or the current frontier for data, multimodal data? Is it synthetic data?

Yann Dubois [27:38] I think synthetic data can probably work well in a data-limited regime. I think multimodal is an interesting one. I definitely cannot talk about what we do internally, but I used to work on multimodal representation learning back in the day, and I always thought that it would really help your reasoning abilities if you have a lot of multimodal data. And I still think this, but, for example, if you look at Anthropic models, they tend to not be that good on multimodal, and they are still really smart.

Yann Dubois [28:21] So it seems that it's not as necessary as at least I would have thought in the past. I still believe that once we go to embodied agents, embodied AI, you will learn a lot about the world, and you will kind of improve general intelligence and usefulness to users by learning how the world interacts with itself. But at least looking, for example, at Anthropic models, it seems that they don't need that much multimodal data to have strong models.

Matt Turck [28:34] And by embodied intelligence, you mean potentially robotics. And so if you use a video that shows how gravity works and how a robot evolves in space, then presumably that would be more useful. Is that the thought?

Yann Dubois [29:04] Yes. The intuition that I think many people had, and I definitely felt for a long time, is that it's hard to understand the world only through text. And it's hard to understand what physics is without really seeing what—for example, you can't understand gravity without really seeing things falling. And when you look at our models, I mean, they kind of understand gravity without having seen that. But it still seems not obvious. It still seems like they would get it more, and they are still missing some common-sense aspects.

Yann Dubois [29:30] So I do feel like we will improve the common sense of our models by having them interact in the real world. But we are still pretty far from that, I think. And by we, I mean just generally the academic community and the AI community seem pretty far from that.

Matt Turck [29:42] And while we're on the topic, as a quick detour that leads us to the concept of world models, taking your OpenAI hat off, are you bullish on world models?

Yann Dubois [29:59] World models in the sense that, yes, you can try to replicate or simulate things, basically work in an environment that is simulated? Yes. The problem is simulations are always going to be really hard and not going to be truthful. So I think there will always need to be a little bit of training that will need to happen in the real world to make sure that the model realizes these mismatches between the simulated world and the real world.

Yann Dubois [30:45] And I think we as a field have a tendency of optimizing something that is simulated or not quite realistic past the point where this is useful. So that's something that I think we should always be careful with: we spend a lot of time and effort on optimizing something simulated and not quite realistic. And it's great at the beginning, but at some point, once you start optimizing too much for something, it's not representative of the real world. And people continue doing that just because that's what they've been doing for a long time.

Demystifying mid-training and post-training

Yann Dubois [31:05] So I just think people need to realize when to stop that. I don't work with these types of synthetic environments as much, just because I don't work on embodied AI. So I don't know if we're at that yet.

Matt Turck [31:18] Okay, great. All right. So, going back to pre-training, mid-training, post-training, let's talk about mid-training. It might be something that people have heard about a bit less. The term comes up a bit less. What is it, and why is it important?

Yann Dubois [31:43] Mid-training. It's just this idea of something that's between pre-training and, as you might realize from the name, the post-training part of the pipeline. And really, the idea is if you have high-quality data that is more representative of what you really want in your final model, you should overtrain on that data. So, taking a step back here, pre-training, what is it? Pre-training is basically trying to learn everything from the world by learning everything from the internet at a high level.

Yann Dubois [32:21] The problem is that most things on the internet are not really useful. If you think, for example, about Wikipedia or GitHub, which is coding data, it just seems like there's way more information in there than some random forums that maybe don't have that much information. For example, ads. There's also lots of ads on the internet. You probably don't want to train too much on that. But in pre-training, we train on everything. And in mid-training, we basically overweight this type of high-quality data that we think is more useful for training the final model.

Yann Dubois [32:44] And this is something—I can't talk about what's happening at OpenAI—but it's something that is definitely happening in all the academic community right now, and all the open-source models have this stage of mid-training.

Matt Turck [32:53] Great. Post-training, let's start at a high level by defining what that is. So, there's reinforcement learning, but that's not the only part of post-training. What else is there?

Yann Dubois [33:17] It kind of depends how you define the term and where you put the boundaries. In my mind, post-training—I'll take it from a very broad sense—includes all the reinforcement learning and the training for our reasoning models. It's just the idea of having something that knows everything about the world and making something that is useful to people. So, pre-training, I think about it—the metaphor that I like giving is you go in the library and you have a lot of books about everything.

Yann Dubois [33:54] And in theory, you can find all the information that you want in the library. But it's much more useful to talk to an expert who has learned these books and who you can ask questions to, and they can answer and understand what you're actually looking for. So this is kind of the goal of post-training at a very high level: making something that is useful to users and is easier to interact with. So there are multiple stages. I'll talk only about things that are happening outside of OpenAI and kind of the usual stages.

Yann Dubois [34:04] There's usually some SFT that is happening.

Matt Turck [34:07] Which is supervised fine-tuning.

Yann Dubois [34:08] Supervised fine-tuning.

Matt Turck [34:08] Yes.

Yann Dubois [34:38] Supervised fine-tuning. And that's actually what, early on, most of the models that were post-trained were only doing: supervised fine-tuning. The idea is that if you have humans that can give you the desired final answer, if you have humans that can give you the gold answer, you can basically clone the behavior of the human. So this is what we call behavior cloning. The problem with this is that you will never get better than what your ground truth gives you. And humans are actually pretty limited in many senses.

Yann Dubois [35:06] So you will never overcome the human labelers that you're working with. The reinforcement learning stage goes from behavior cloning to really optimizing rewards. So the idea is, I don't know what the ground truth is. I don't know what the perfect answer is, but here's how I would say whether the answer is correct or not. And here are the things that I want in the answer. And what you do is you start optimizing. You start having a model that tries to get more reward, basically optimize this reward function. That's what we call it. And it goes beyond what you currently have, what humans can do, or at least what the humans that you're working with can do.

Yann Dubois [35:42] So this, I would say, is the two big stages. Then in reinforcement learning, that depends on which models are being trained. At least in the open-source community, it seems that there are different ways of doing that. Reinforcement learning when you have verifiable rewards, so reinforcement learning where it's really easy to say whether something is correct or not, and you can really kind of have a binary reward for this.

Yann Dubois [36:11] And that goes back to how we talked about o1 and o1-preview in the past. And then you have reinforcement learning without verifiable rewards, where maybe I could do pairwise comparisons. I can say this answer is better than this other one, but I don't really know. I cannot quite say this is the perfect answer. So, of course, it's a continuum and there's everything in between, but I would say these are the three high-level things to think about when you think about post-training in general and how people are usually doing it in the open-source world, is that they take SFT, they clone the behavior that you can collect online or from humans.

Yann Dubois [37:04] And then once it's already at a pretty good level, they just do this reinforcement learning to go beyond what we currently have. Because if you just started from reinforcement learning, it would be very inefficient. Because the problem with reinforcement learning is that you have to stumble across the right answer, basically. Because how reinforcement learning works is you sample many times, essentially, from the model that you're training, and you say, this one is correct, this one is not. And you say, do more of the one that is correct.

Does RL create new capabilities in AI?

Yann Dubois [37:21] So you have to stumble across the right solution. So you're much better off first getting as close as possible to the best you can do, and this is behavior cloning, and then doing reinforcement learning.

Matt Turck [37:29] Does reinforcement learning create new capabilities, or does it make the model better at existing capabilities?

Yann Dubois [38:04] It's really hard to say because pre-training, when it's trained on all of the internet, arguably already has all capabilities in it. So it would be even harder to answer this question scientifically because arguably everything is already there. What I would say is that if you look at models that we were training or post-training, like two years ago in the open-source world—for example, I worked on one of them, Alpaca, where we used 50,000 examples for SFT—and now when you look at reinforcement learning from models like Kimi or from DeepSeek models, it seems that they are closer to one million data points.

Yann Dubois [38:50] So definitely people scaled up the reinforcement learning stage a lot. And from this, it seems that they've learned new capabilities, like this reasoning aspect, this fact that you can check your answer and try to improve it. So you can really think for longer to get a more correct answer. So all this to say that arguably everything is already in pre-training, but we were definitely able in the last year and a half, even in the open-source world, to have more capabilities.

The challenges and frontier of scaling RL

Matt Turck [39:20] I heard several times that reinforcement learning is pretty finicky and hard to scale. And part of the reason why we as an industry didn't do reinforcement learning as part of the initial kind of LLM progress curve was precisely that, that it was hard to make work. What is hard about scaling RL? Is that a question of datasets, knowing where the rewards are, or something else? Yeah, I think it's a combination of things.

Yann Dubois [39:46] I would say most people who did not work in reinforcement learning in the academic and research community up to two years ago probably thought reinforcement learning just doesn't work and is too finicky to work with. I used to be that type of person. And actually, when I saw ChatGPT come out, they had this blog. I was not at OpenAI at the time. I saw this blog that says that they use reinforcement learning. And my first thought was, I can do the same without reinforcement learning.

Yann Dubois [40:17] Because this is just an overcomplicated method. And this is actually the project that we started working on with Alpaca, was exactly, let's try to reproduce that only using SFT, just by doing this behavior cloning. And, for example, Yann LeCun famously gives this metaphor that reinforcement learning is just like the cherry on the top. So I think that was really the intuition that most people had. It seems that after crossing a certain scale of models that know basically everything about the world and what we call good priors about the world, it seems that reinforcement learning just started to work.

Yann Dubois [40:52] And this is not only with LLMs. Robotics seems to be entering the same stage where they're realizing that actually it used to be very finicky, but now that we use models that know already everything about the world, it actually learns pretty well. Now, to answer your question about what is still complicated with reinforcement learning, one is an infra aspect. So, just systems in general. Reinforcement learning, you have, at a very high level, basically to sample, as I said before, many answers and say what is correct and what is not.

Yann Dubois [41:30] And this sampling is just very expensive, and you have to do it at scale. The other issue that also in the open-source world people are seeing right now is that when we are training more agentic systems, you only know whether you're correct at the end of your very long rollout. So you get very little information per token of whether you were correct or not. And it's hard to basically do attribution. It's hard to say what part of your entire answer was the one that led you to being correct.

Yann Dubois [42:01] So that's more of an issue on the machine learning side. The ideal world in machine learning is when I can say exactly, like, this thing was good, do more of that. And the problem again with these agentic systems and reinforcement learning with these agentic systems is that you don't really know which part was good or not until you arrive at the end. That's another big issue for reinforcement learning.

Matt Turck [42:17] What's the current frontier of reinforcement learning? It seems like there's a jungle of acronyms like GRPO and other techniques. What are you using? What are you excited about? What do you think is promising?

Yann Dubois [42:47] I can't talk about what we're using, but, for example, in the open-source world, GRPO seems to be working very well. And people used to have different methods like PPO and DPO, and people seem to have really converged to this one. The big difference with other methods is that you, again, do this simple method that I told you about: sampling as many answers as possible, and you say which one is correct. So in some way, GRPO is a very simplistic method.

Is building AI models a craft or a strict science?

Yann Dubois [43:09] And in general, we saw over and over again in machine learning that the simplest method where you can scale up in terms of compute is usually the one that ends up working the best. And that is kind of what is happening here, at least in the open-source world.

Matt Turck [43:27] As you described some of the challenges, a question crossed my mind. You often hear that AI systems are not built; they're grown. How would you characterize it? What part is science versus craft, or trying multiple things and then just keeping what works best in your day-to-day life?

Yann Dubois [43:54] Yeah, that's a great question. I think how it usually works is that it starts being craft. People just try out many things, and they start building a mental model of what works and what doesn't. And over time, we move from this craft land to more science. Science, or more scientific approaches, are really the ones that first end up working. It's very rare that you take a really scientific approach and you say, "This is the optimal thing to do," and you do it and it just works.

Yann Dubois [44:42] There's some sense of alchemy. People just have a good flair for something, and they make it work. And then other people, or that person, start trying to improve what we are doing by being very scientific. And I would say this happens over and over in machine learning. So first craft, then science, and both are really important, but it's different stages of the pipeline. In terms of engineering, this is definitely something that is always necessary. So I would say most researchers have moved to being relatively good at figuring—at least, I wouldn't say good engineers, but good at working in complex systems and figuring out what they need to try out.

Yann Dubois [45:09] And the systems and the infra that we have have become more and more complicated. So definitely the work required changed over time.

Matt Turck [45:42] Fascinating. All right, so still in reinforcement learning and circling back to some of the things you said at the beginning: if I want to make my model better at computer use or genetic coding or whatever domain, then I would spend a particular amount of time doing specifically reinforcement learning for computer use and putting together a dataset and then coming up with rewards. Is that how it works? You just pick one problem and you just do reinforcement learning specifically for it?

Yann Dubois [46:10] To be clear, I talk more about reinforcement learning because this is the part I know the best, and this is what I've worked on for a long time. We talked about mid-training before. All these things are also extremely important, and you can improve it in different parts of the pipeline. As I said before, the closer you are to the final stage of the model, usually the smaller the scale of the training becomes. So you can iterate fast on that because now you can iterate in terms of days rather than months.

Yann Dubois [46:46] So usually people start from this fast iteration loop, and then they go deeper and they make bigger changes across the entire stack. So this is not to say that only reinforcement learning matters. I'm really not saying that, but it's just that that's where people will start doing changes, and then that will permeate and we will go deeper into the stack. So this is how it works. And in the open-source world, it's very much like that too. I think you see way more post-trained models than you see new pre-trained bases.

Yann Dubois [47:06] And you see way more improvements in the algorithm. And that's why we talked about GRPO, DPO, PPO. There are so many XPOs, and that's because people can iterate really quickly on this final stage of the pipeline.

Matt Turck [47:26] And the jagged nature of those models, does that come from this approach of picking this problem and that problem, and therefore it's going to be excellent at those problems but not as good at other problems? Or is that a more fundamental characteristic? There's definitely some of that.

Yann Dubois [47:53] For sure, if you optimize more on specific types of problems, you'll be better in that setting. I would say, at least my intuition is that it's less about the exact problems that you're optimizing on, and it's more about the class of problems that you're optimizing on. So, for example, if you are really good at math competitions, your model will probably be pretty good at coding competitions. So it's not about the domain; it's more about the skills that are necessary and the way to think and the horizontal capabilities that you need for performing these tasks.

How AI models generalize across different domains

Yann Dubois [48:21] And that's what I think you're usually seeing when some model is really bad at something: it's actually bad at that in any domain, in any language. So you have to think about this domain and then this generalization of this domain, not necessarily per-domain capability.

Matt Turck [48:57] So speaking of generalization, there's been that clear evolution from math and coding success to now starting to cover different areas. So that's the whole GDPval thing, where across the economy, different areas are being evaluated in terms of model performance. Sort of same question: is that the result of overall model progress? Or is that a deliberate, okay, now we're going to take this part of the economy and build a dataset for it and do mid-training and do post-training? How does that progress work from those very specific domains to generalizing to the rest of the world?

Yann Dubois [49:33] It's definitely something that we actively push on. I think people are realizing, us and also other companies, that we are moving towards this world where we want to really make products that are useful and improve productivity of people and help people in their day-to-day life. So I think there's a very active move to deciding what are the domains that we should be prioritizing. Now that we know we have an algorithm that we can apply in different places, what we are constrained by is more collecting the right data, having people who really care about a certain problem work on that problem.

Yann Dubois [50:20] But there are not that many people who can do these things. So you really need to prioritize. So this is a very active, proactive kind of approach here. And in general, I would say the performance of the model really depends on the number of people who care about the final output of the model and who are looking at that model. So if they start looking more at specific verticals, these verticals will improve really quickly. But again, we don't have that many of these people that can do these things.

Matt Turck [50:53] But to unpack something that you alluded to, I think a minute ago, do models actually generalize now more, especially from a reinforcement learning perspective? So making a model very good at domain A or B is likely to make the model better at C, regardless of the amount of effort you put into developing rewards for domain C?

Yann Dubois [51:19] So I think there are different axes of generalization. One, there's algorithmic generalization. And that's really, can I use the algorithm that I developed or this black box that I developed for domain A, and can I use it for domain B? And again, even talking about the open-source world, it really seems that people are able to do that. They take GRPO, they apply it in many different places, and it just works. So that generalization seems to be relatively good, which is why we are seeing a lot of progress.

Yann Dubois [51:47] Otherwise, it would be hard to make progress. Then there's the generalization of the model that is trained on one particular dataset. And this is what I was alluding to before, is at least my mental model is that generalization happens in terms of capability. Like if the capability is the same, you will see generalization across domains. Again, like different languages, like coding, like you can optimize for C++ coding, for having a good C++ model with very little RL in C++, partly because this pre-trained model has seen all of C++.

Yann Dubois [52:38] And so it already kind of understands the basics of that language. So that type of generalization definitely happens. The generalization that I think is harder is when we don't have these horizontal capabilities. So I'll give you one concrete example. If my model is very intelligent in terms of being correct on competitions, I usually take that example because it's somewhat contrived, at math competitions, coding competitions. From a human perspective, people that are good at these things are usually just smart. And if they are smart, or someone might think that at least, that they're just smart.

Yann Dubois [53:05] And if they're smart, they can actually do other things too. But that is really not true. And that type of generalization is really not true, because many things where we need to have humans working on expert domains, the world is very messy, and these coding competitions and math competitions are extremely well specified. And you need to have the capability of understanding underspecified tasks, understanding how to deal with the messy world, and understanding what are even the resources that you need to answer the question.

Yann Dubois [53:45] If you look at the math competition, you usually have everything in the prompt. It's like you have five lines or maybe 15 lines, and it's all the information that you need to answer this question. In the real world, if I'm a consultant, if I work in finance, I need to go on the internet. I need to find and extract different information just to understand, before doing any of the reasoning, just to be able to do that reasoning. And this type of horizontal capability is the thing that doesn't usually generalize. If you have that horizontal capability, but in many cases we don't have that horizontal capability.

How reinforcement learning cures AI hallucinations

Yann Dubois [54:18] So yeah, that's why we hallucinate, actually, in every domain. When you have hallucination in LLMs, if a model is really bad at saying that it doesn't know, that usually happens in every single domain. You won't have one domain where the model is extremely calibrated about its knowledge and another domain where it's not.

Matt Turck [54:30] And as a quick detour, is hallucination also a reinforcement learning problem where you reward the behavior to say, "I don't know," when it occurs?

Yann Dubois [55:00] John Schulman has a great presentation about that, I think from one or two years ago, where he was saying that if you do behavior cloning, so this SFT that we talked about before, you will basically reward and optimize for hallucination. If your model doesn't know about something, but now you say that the right answer is to say something. So I'll be very concrete. If the model doesn't know about a paper, and now in an answer that is given by a human, you say, "Here's where I got the information."

Yann Dubois [55:40] And then you cite that paper. What you're actually optimizing the model to do is cite something that doesn't exist, because it doesn't know that that paper exists. And so John Schulman had this great presentation saying SFT is going to force hallucination. While in reinforcement learning, given that, as I said, you sample from the model in the first place, it's extremely unlikely that it samples something that it doesn't know and it's correct. That's extremely unlikely. So you will never reward that behavior.

Negative generalization and conflicting instructions

Yann Dubois [56:04] You will only sample things that it doesn't know and is being incorrect, and then you will kill that behavior. So hallucination, at least the intuition that people have, is that it can come, for example, from SFT, and it can come from this post-training pipeline. But if you have a good reinforcement learning pipeline, that shouldn't happen too often.

Matt Turck [56:24] Going back to generalization as well, are there examples where actually getting better at one domain makes the model worse at the rest? A little bit to what you were saying about some people are very good at math, some people are very good at English. Pretty often, they're not the same people.

Yann Dubois [56:52] In domains, usually not. What will happen, though, is you will make decisions based on which domain we optimize for. And if you optimize for one domain, you will be able to optimize less for another one. So it's not necessarily that optimizing for one thing will make the other one worse. It's just that, as a result, you can optimize less for the other one because you're compute-constrained, you're data-constrained. You have your human bottleneck also in terms of that work.

Yann Dubois [57:25] What does happen is you can have negative kind of generalization, like bad generalization or negative transfer, more for these horizontal aspects of the model. So I'll give you a very concrete example: explicit instruction following versus implicit instruction following. If I have a model—and this is, we often hear, for example, from OpenAI models—that they tend to be really good if you tell them exactly what you want. But as a result, sometimes we hear also that they're less good if you are not as specific about what you wanted.

Can RL scale to law, medicine, and the broader economy?

Yann Dubois [58:05] For example, if I make a typo and I say, change this file, and I make a typo in this file, an extremely good model at explicit instruction following will change the wrong file, the one that has a typo. But humans would probably realize that you made a typo. And as a result, there are cases where this explicit instruction following goes against this implicit instruction following. So you will have cases where basically these horizontal capabilities go against each other.

Matt Turck [58:27] And maybe to close on this whole reinforcement learning: is your sense that, as we progress from being excellent at coding and excellent at math and move to the rest of the economy, do you think that the rest of the economy is a tractable problem? Do you think we can get to the same level of performance ultimately?

Yann Dubois [58:54] Yes, we can. I don't think there's anything really deeply special about these domains where we cannot optimize and where we couldn't get the same with other domains. The but is for at least two reasons. The first one is most of the people working on these models are pretty good at coding, and they really care about coding because that's what they use as their day-to-day kind of drivers. And there's nothing better than the user being also the one who trains the model, because then they understand the issues.

Yann Dubois [59:33] It's very hard to really, for me, for example, it's very hard to really understand what should we change on the legal aspects of the model if I don't understand anything about the legal domain. So that's one thing. The other thing that you will often hear about, and I mentioned also briefly before, is this kind of verifiable rewards. There are domains where it's easier to say whether something is correct or not. For example, in the case of cyber, you mentioned that before, that cyber has been improving a lot, the cyber capabilities of our models.

Yann Dubois [1:00:09] And this is because in cyber, it's extremely easy to say, are you correct? Did you find—did the cyber issue that you find is a real issue or not? It's very easy to test it. And so there are domains where reinforcement learning is just easier to apply, but there's nothing, I would say, in the capacity of the model that is constraining the model to be as good at legal and medical and other domains. So, the short answer is we know less about these domains, and definitely there are some domains that are easier to optimize for in reinforcement learning.

The evaluation bottleneck and Model as a Judge

Matt Turck [1:00:30] Great. Let's talk about evals for a minute. That's a hugely important topic. Maybe to start, why is it so hard to evaluate a model in the first place?

Yann Dubois [1:00:51] Evaluation has been harder and harder as models become better. And that's because the tasks that we ask of the model become more and more general and more and more open-ended. So now I maybe just say, "Build me a website that does X." Well, before, in the past, I would just be like, "Hey, is there a specific bug in this implementation that you have?" And it's much easier to say whether there's a bug because I can inspect, I can know, I can have a human that says, "Here are all the bugs that you have," and then you can apply that automatically.

Yann Dubois [1:01:33] While the website one is very hard to know what is the optimal answer, because there are many good answers. There are many good ways of building a certain website. This open-ended nature of models really makes evals harder. There's also another issue: models, in specific axes, are becoming better than the majority of humans. And so we have fewer and fewer humans that can actually evaluate these models in particular axes. So that's definitely a constraint. Another one, to be honest, is kind of cultural.

Yann Dubois [1:01:58] Most people want to improve the model, and they think that the best way to do that is kind of training the model, when in reality, finding issues and making sure that we can quantify improvements is just as important, if not more important. But there's always this cultural gap. That was especially true, I would say, in the academic world up to two years ago, when evals were always fixed, benchmarks were always fixed, and even datasets were kind of always fixed, maybe, let's say, four years ago.

Yann Dubois [1:02:27] And there was a mentality shift of, like, okay, data is actually critical. And now there's a lot of people working on data. And I think evals were still not quite there. People don't really fully—everyone knows that it's important, but people don't really understand how impactful it could be to work on evals. So actually, my first project at OpenAI, I just came in and I was like, "I want to work on data and evals because I know that this is the thing that no one is working on."

Yann Dubois [1:02:45] And as a result, I know that it's super impactful to work on that. And yeah, the tide is shifting, but not fast enough.

Matt Turck [1:02:59] And is the pace of progress in model as a judge and AI evaluating AI, is that moving as fast? Is that a distinct part of research, or is that fundamentally the same?

Yann Dubois [1:03:26] It's really fundamentally the same method. Most of the things that we do in evals, especially now that we have reinforcement learning, could just be applied nearly exactly as is during training. So that's another reason, actually, why evals are so complicated, is that every time you build an eval, you actually build a way to build training datasets. So now you're going to optimize that training dataset. Well, even if it's not that eval, it's going to be the same type of data, and now you're going to do super well because we have this generalization of capabilities that I was telling you about.

Yann Dubois [1:04:07] You will learn that on that other dataset, and now you'll become really good at that eval, and that eval will become obsolete really quickly. So that's also an issue with evals. But yeah, to come back to your question, model as a judge, it's really important. And I think it's probably one of the most important things because as we get better models, we have this self-reinforcing loop and we have this capability flywheel where better models become better teachers for other models. And this is really important for training, but then you can also do the same thing for evaluation.

Continuous AI progress & continual learning

Yann Dubois [1:04:21] So a lot of my team works on that, and I think what's really critical is to work on this model-as-a-judge kind of framework.

Matt Turck [1:04:49] Okay, fantastic. All right, so as we get towards the end of this conversation, I'd love to zoom out a bit and get your sense for where things might be heading. Obviously, it's incredibly hard to make predictions on AI years out, but let's call it the next 12, 18, maybe 24 months. Is your sense that things are going to continue progressing, or are we heading towards something that could feel more like a discontinuity?

Yann Dubois [1:05:19] In terms of progress, as I was saying before, I think it's always continuous. Now, the feeling of discontinuity will happen. It did happen three months ago with coding, or four months ago with coding. And I think that will happen now in every other domain. Most people are not feeling the capability and usefulness of our models the same way coding and software engineering are feeling right now. So this will definitely permeate, I think, through many other verticals.

Yann Dubois [1:05:48] Now, in terms of capability bumps, in terms of, let's say, the verticals that we're already looking at, I think it'll be more continuous, and there will never be big discontinuities. Most of them are always local discontinuities, but you zoom out and it always just feels pretty smooth. It's not always like this, but that has been the case most of the time, and I definitely cannot predict when is the next big discontinuity.

Matt Turck [1:06:09] What is your sentiment on this general concept of accelerating loops in AI? So whether that's continual learning to make models more current and able to learn faster, to this broader concept of AI building AI in an increasingly automated way, fact versus fiction, and what are you excited about?

Yann Dubois [1:06:35] I'm extremely excited about continual learning. I think we haven't quite cracked it. I mean, we have Codex Memories, and that is helpful, but it's definitely not the end state. I have a friend who always tells me about, again, another type of plot that we should be looking at, which is: x-axis, time; y-axis, utility that you provide to users, or usefulness, basically, of the models. And right now, actually, most models at day zero, if you just drop them in a company, arguably they're more useful than most new employees.

Yann Dubois [1:07:12] So they start higher at T-zero, but then across time they're mostly constant because they don't really learn company knowledge. They don't really learn to be more efficient over time at doing the things that they are doing, while humans learn really quickly. And what is important is this integral, or the area under the curve of these curves. And as a result, I think humans are still more useful in many cases. And that's why what we will need is continual learning, to make this curve monotonically increasing over time and basically make models more and more useful the longer they work in a certain environment.

Yann Dubois [1:07:50] So I'm extremely excited about it. I'm actually surprised that we're not quite there yet. Three years ago, when ChatGPT came out, I remember I was doing a startup with friends, and we were thinking about working on continual learning and personalization and memories in general. And we were like, "Ah, OpenAI is going to do that in the next six months. They have all the data. They're going to figure it out, and they have all the users, and the models are going to learn super quickly from users."

Yann Dubois [1:08:01] And three years later, I don't think we're there yet.

Matt Turck [1:08:06] And quickly, in layman's terms, what is the fundamental difficulty?

Yann Dubois [1:08:31] It's a good question. I actually don't quite know, to be completely honest with you. I don't quite know why it's taking us that long to figure it out. It's this type of domain that I think, if we really put enough resources behind it, we would figure it out. Of course, especially when we talk about this memory inside of a company, there are big questions about permissions. And there are a lot of questions about privacy and what you can share and what you cannot across models, across users.

Will foundation models eat the agent harness?

Yann Dubois [1:08:50] But for a single user, even for a single user, we're not quite there. And I don't quite know why. At least at the high level that I can talk about, I don't know why.

Matt Turck [1:09:21] Yeah. What you bring up is, I think, really interesting for AI builders and investors and startups, which is this question of the models getting increasingly smarter within an enterprise. And in particular, there's this whole tension between what the models are able to do and what a lot of people have built around the model. So a year or two ago, it was RAG. These days, it's all about harnesses for agents. And a lot of people are wondering whether the models are going to end up eating the harness, whether the harness is just a temporary thing.

Matt Turck [1:09:35] From your perspective, what do you think happens?

Yann Dubois [1:10:07] Yeah, I think harnesses can really improve the capability of a model right now. I think, given that we are seeing this really fast progress in terms of capability, I personally wouldn't push that much on the harness unless the harness is for a very concrete goal that you're trying to achieve right now. So certain companies, if they are focused on a specific vertical, want to go from this 80% reliability to maybe 85%. Harnesses will give them that.

Yann Dubois [1:10:39] And I think that's very important, but they need to do it while knowing that they will have to retune that harness in the future. And I think that's totally fine. If you try to have a general harness that will sustain over time, I don't think that will work. Harnesses for specific domains are a short-term thing that you need to do. I think there will always be so much you can do in harnesses. And if anything, I think everyone should do more of that if they have a specific problem in mind, because we're leaving so much on the table without a good harness.

Yann Dubois [1:11:09] Arguably, I think if we froze the models that we have right now and you really worked on the harness, and maybe we also spent more time training with a great harness, I think people would really feel the AGI in every single domain, or could already feel that in every single domain. But given that we're not freezing it and we're going to continue training better and better models, I think we don't really understand what the final harness will be, and it will always change.

Why startups should focus on the last mile of AI

Matt Turck [1:11:50] Same question about applications. So we alluded to your progress in different verticals: 1% on Office QA Pro. So, bit by bit, you're doing more and more of this. So do you think people should be building applications anymore? Or is, ultimately, as we get closer to AGI, all of this going to be part of the model capabilities?

Yann Dubois [1:12:35] There's so much space for external companies or startups pushing on specific verticals. I think there's so much space for that. The reason why is because a lot of people kind of think about intelligence, in quotations, and raw capability as being the real bottleneck, but I don't think that's true. I think most of the time, the bottleneck is the last mile. It's making sure that the model has access to the right permissions or also has access to the right connectors and things like this.

Yann Dubois [1:13:09] And we are going to be very focused on this general aspect. And I think there are other companies that should be focused more on the verticals and providing maximum value of what we currently have. So I think there will always be a lot of space left for this last mile in different verticals. And I would highly encourage people to continue working on that. And maybe one day, when we stop making horizontal progress, which I don't think is anytime soon, maybe we will start focusing on that.

Yann Dubois [1:13:21] But yeah, that's not what we're doing now.

Matt Turck [1:13:33] Okay, well, that feels like a very optimistic note, at least for the startup ecosystem, to end up on. Thank you so much, Yann. This was terrific. Really enjoyed it. Thank you so much for spending time with us.

Yann Dubois [1:13:35] Great. Thanks, Matt.

Matt Turck [1:13:56] Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.