Sonnet 4.5 & the AI Plateau Myth — Sholto Douglas (Anthropic)
The MAD Podcast with Matt Turck · with Sholto Douglas, AI researcher, Anthropic
Sholto Douglas is the AI researcher at Anthropic. We cover why Sonnet can surpass an earlier Opus because mid-tier models iterate faster, how reinforcement learning and test-time compute create a new scaling axis for reasoning, and why 30-hour coding agents depend on memory and self-correction rather than raw programming ability.
Chapters
- 1:09 — The Rapid Pace of AI Releases at Anthropic
- 2:49 — Understanding Opus, Sonnet, and Haiku Model Tiers
- 4:14 — Shelto's Journey: From Australian Fencer to AI Researcher
- 12:01 — The Growing Pool of AI Talent
- 16:16 — Breaking Into AI Research Without Traditional Credentials
- 18:29 — What "Taste" Means in AI Research
- 23:05 — Moving to Google and Building Gemini's Inference Stack
- 25:08 — How Anthropic Differs from Other AI Labs
- 31:46 — Why Anthropic Is Laser-Focused on Coding
- 36:40 — Inside a 30-Hour Autonomous Coding Session
- 38:41 — Examples of What AI Can Build in 30 Hours
- 43:13 — The Breakthroughs That Enabled 30-Hour Runs
- 46:28 — What's Actually Driving the Performance Gains
- 47:42 — Pre-Training vs. Reinforcement Learning Explained
- 52:11 — Test-Time Compute and the New Scaling Paradigm
- 55:55 — Why RL on LLMs Finally Started Working
- 59:38 — Are We on Track to AGI?
- 1:02:05 — Why the "Plateau" Narrative Is Wrong
- 1:03:41 — Sonnet's Performance Across Economic Sectors
- 1:05:47 — Preparing for a World of 10–100x Individual Leverage
Transcript
The Rapid Pace of AI Releases at Anthropic
Matt Turck [1:00] Sholto, welcome.
Sholto Douglas [1:10] How you doing? Great to be here.
Matt Turck [1:30] Sonnet 4.5, which is the big news of this week. 3.7, which was like this huge deal at the time. In my mind, if you had asked me, I would have said, “Oh no, that was last year.” But in fact, it was just in February of this year. What's the right way to think about that pace of releases? Is that a proxy for progress accelerating?
Sholto Douglas [2:11] Yeah, I think it's a proxy for a couple of things. One is that there's now this two-paradigm regime where previously you did pre-training scaling and reinforcement learning scaling, and now we're in a mix of the two, basically. And so I think that gives you more opportunities to update models, because it means that you can make advancements along multiple frontiers. And then that means that you sort of end up shipping more frequently. We're now 2.5 years after ChatGPT. And so the post-ChatGPT investment cycle is finally hitting where compute availability is increasing and all of this.
Sholto Douglas [2:42] And so it means that you should expect, actually, the pace of progress to increase, because there's lead times in commissioning chips, basically. So even if you wanted chips last year, it would have been impossible to get them because TSMC was booked out and so forth. So finally, this year is where the compute supercycle is beginning properly, in effect. Yeah.
Understanding Opus, Sonnet, and Haiku Model Tiers
Matt Turck [3:00] Okay, great. Maybe for situational awareness for people listening to this: Sonnet, Opus, is there still Haiku somewhere? Maybe walk us through the differences between those.
Sholto Douglas [3:30] So we release models along three categories, three tiers: Opus, which is the smartest model; Sonnet, which is the mid-tier model; and Haiku, which is the fastest, cheapest model. One of the interesting things about this most recent release is actually Sonnet is smarter than Opus. And this has happened before. In fact, this happened last year. It's a reflection of fast progress because it is cheaper to train mid-tier models than large models. And so what happens is that you end up doing a lot of progress on smaller models.
Sholto Douglas [4:06] Eventually, you need to choose when to scale up and sort of get the benefits of scale in a model. Often, you make progress fast enough that your mid-tier model is great anyway, and it's actually better than the large scale-up model that you did previously. And I think this is also a little bit of a reflection of the reinforcement learning paradigm, where you can take a model and you can train it and extend it with reinforcement learning, basically. So that allows you to take a mid-tier model and make it as good as a larger-tier model of six months ago or three months ago.
Shelto's Journey: From Australian Fencer to AI Researcher
Matt Turck [4:29] All right, so before we go into all of this in greater detail, I was curious about your story, your journey to Anthropic, and then what you currently do at Anthropic, how you would describe your role.
Sholto Douglas [4:34] Yeah. So I think, how far back do you want me to start? From the beginning.
Matt Turck [4:35] From the beginning?
Sholto Douglas [5:02] Yeah. So a couple of things. One is that growing up in Australia, there's a very traditional set of paths you can take. You can become a lawyer, you can become a doctor, or you can go into finance. Australia is a wonderful country in so many ways. In particular, the quality of life is so high that it means that people just choose these default paths, have a fantastic life. And I was very lucky in some ways. My mum was actually frustrated in her ambitions.
Sholto Douglas [5:29] And so this meant that I had the perfect mentor throughout my entire life. She studied medicine, went on to do emergency medicine in South Africa, but wasn't ever quite able to break into public health in the way that she wanted to. She wanted to do systemic change in public health. And at the time, that was just very difficult for a woman. So instead, I had her full attention. Growing up, when I did an exchange in China, I got this dossier this thick of China's political economy and different actors in the current startup ecosystem and this kind of stuff.
Sholto Douglas [6:08] So I had this wonderful, constantly driving education in a really supportive and wonderful way. I also was lucky enough to get into fencing. And through fencing, I had the experience of becoming one of the best in the world at something via repeated effort. I became top 50 in the world, at my best, 43rd. And it was partially a consequence of—well, I think in large part due to having a coach and perfect mentorship that was one of the best in the world.
Sholto Douglas [6:38] He moved to Australia because his wife was Romanian. He had just coached Italy to the gold medal in the Olympics. She was facing discrimination in Italy. And so I had, on the one hand, perfect academic mentorship, and on the other hand, perfect athletic mentorship and a proving ground to grow up watching these people on YouTube and then become one of the best in the world at something.
Matt Turck [6:41] Early introduction to reinforcement learning. Like, do this, don't do that.
Sholto Douglas [7:00] In some ways, yes. Or in other ways, like an introduction to, you can watch these people on YouTube and analyze what they are doing to become who they are, and replicate that. And you could be part of that world. All it just takes is intense amounts of effort.
Matt Turck [7:12] It's a theme that I find fascinating, that across any field, the fundamental impact of YouTube and the fact that regardless of the field you look at, every kid seems to be just much better than the prior generation.
Sholto Douglas [7:13] Right, exactly.
Matt Turck [7:14] I don't know if it's been studied, but at least that's your experience.
Sholto Douglas [7:41] Yeah. And I mean, I think we should see the same thing with AI, right? Like, in the same respect, everyone will now get a perfect tutor. I then actually had that experience again with AI. Fencing wasn't something I wanted to do ultra-long-term. I wanted to take a shot at the Olympics and then try and progress into working in technology, basically. I was very lucky to read a Gwern essay on scaling, where he basically details the scaling hypothesis. After reading that, I was like, oh my God, this is absolutely, clearly, AGI progress over the next decade is going to be one of the most meaningful things to work on in the world.
Sholto Douglas [8:01] It's the largest lever we have to meaningfully advance the world. And so I started doing my own research on nights and weekends, and I was, like, partway through undergrad.
Matt Turck [8:01] How old were you then?
Sholto Douglas [8:06] So this is last year of undergrad and the year after.
Matt Turck [8:08] And in undergrad you did?
Sholto Douglas [8:27] I did computer science, robotics. And I sort of vaguely grew up looking up to Elon Musk and this kind of stuff. I wanted to build rockets and Tesla, but I didn't have a concrete idea of what actual problem I wanted to solve. Reading that essay was the critical hinge of, okay, AGI is possible this decade. It seems like the most meaningful thing in the world to work on, and I need to figure out how I can demonstrate that I should be working on this.
Sholto Douglas [9:02] I was at the time working on robotic manipulation stuff. And so I started working on scaling up robotic manipulation, trying to train general foundation models for robotics from the bedroom, which is now a big thing. There's a lot of general foundation model for robotics companies. It was a little bit early then, but I rigged up my own simulator, collected a lot of teleoperation data, trained models, got a loan of TPUs from Google. Eventually, some people at Google noticed the work I was doing and said, hey, this is great work.
Sholto Douglas [9:28] Would you like to come work with us? It was actually very fortuitous because, for example, I didn't get into the PhD programs that I wanted to. I applied to a couple of PhD programs here after undergrad and didn't get in. But I was very lucky that the work that I was doing really resonated with Google. And so they reached out.
Matt Turck [9:47] Which is a fascinating concept, that at some point you could have had an academic roadblock, but still succeed to the extent that you're currently succeeding. For people maybe outside of the AI research world, it sort of feels like whoever is the smartest academically wins.
Sholto Douglas [9:47] Right.
Matt Turck [9:58] But does that suggest that being great academically and being a great Anthropic researcher are two different things? You need slightly different qualities?
Sholto Douglas [10:15] I think they're very highly correlated, but I think the signals that are usually used to gate academia are, like, there are dramatically more people that satisfy the criteria of being really effective than there are that have the correct signals that would then enable them to progress to the next stage in an academic career. For example, if you're here in the US, you end up doing, as an undergrad, research that can get you a NeurIPS or ICLR paper, whereas in Australia, it just isn't the case.
Sholto Douglas [10:50] Right. I remember Pieter Abbeel actually once visited our lab in Australia and asked people to put their hands up if they were going to NeurIPS, and no one put their hands up, not even the PhD students. So it means you don't have, again, that mentorship aspect that is so important. And so you don't get a chance to develop problem taste on the things that mattered, and therefore you don't have the correct signals that indicate you would have high potential for academia. I actually think that right now a lot of the signals we look for aren't traditional PhDs or anything like this.
Sholto Douglas [11:17] I mean, this is obviously very useful, but the fastest route, or the most immediate one, is whenever we see a really good blog post where people have done an incredible amount of work in an independent fashion, it's one of the highest-signal things there is. One of the examples I love to use here is this guy called Simon Boehm, who's one of the leads on the performance team at Anthropic. And he's published, to date, the best guide on how to optimize a CUDA matmul on a GPU.
Sholto Douglas [11:33] It is simply the world's best CUDA matmul guide. No one has done this for attention, right? If someone did this for attention, then, I mean, we would reach out with a job interview offer the next day, right?
Matt Turck [11:34] Yeah.
The Growing Pool of AI Talent
Sholto Douglas [12:01] And in fact, someone did it for TPU and for some kind of attention the other day. And we were like, this guy, immediately, let's send out a request to interview. So I think there is actually an absence of agency, there's an absence of taste, and there are still many ways to demonstrate this, usually by producing a world-class artifact in some independent fashion.
Matt Turck [12:16] Yeah. And a little bit to this conversation and the YouTube discussion, do you see the pool of talent in AI research, whether academically sanctioned or sort of more indie, is that growing?
Sholto Douglas [12:36] Yeah, I think it's growing quite dramatically. I also think we have done a pretty good job of growing people. And I mean, I think Anthropic has taken many, many junior people and grown them into really fantastic researchers and engineers in quite a deliberate way. So I think it's definitely growing.
Matt Turck [12:40] So Google noticed you, and then what happened?
Sholto Douglas [13:13] So Google noticed me. I started at Google, I think, like a month before ChatGPT or something like this. So it was actually a fantastic time to start at Google, because the entire company was suddenly forced to react instantaneously and compete with OpenAI. So it meant that there was this gap of, I guess, the typical command structures and everything were not well suited for that particular battle. It wasn't a preexisting org. Gemini was sort of forged out of the merging of Brain and DeepMind.
Sholto Douglas [13:47] It meant that there was just a huge gap in terms of agency, really, of figuring out what we needed to do, doing it as fast as possible, organizing people together to work on important things. And so I ended up, one, getting the chance to develop a lot of taste by working closely with people in those early months of Gemini, but two, also quickly got the opportunity to step up and lead various parts of this. So one example of this is we just didn't have an inference stack that was at all sensible for the modern world of LLMs.
Sholto Douglas [14:27] And so we had to notice that, design one from scratch. A lot of the things you now see in the sort of SGLangs and stuff of the world are things that we had to derive from first principles at that point in time. And we wrote the inference stack. This ended up saving several hundred million dollars, I think, even over the first six months, and meant that I was then trusted to solve both hard technical and sociopolitical problems. Because one of the interesting things about the inference stack as a problem was that it was both a really large technical challenge and a large sociopolitical one, because the ownership of the preexisting stack was distributed across five or six different teams.
Sholto Douglas [15:02] And it meant that actually enacting change was quite hard. And so it was a challenge along multiple dimensions. That then led to me having the trust to solve problems of this form across other parts of the ML stack. And so, for example, later on, when the thinking stack was started, the reasoning stack, I was in charge of research infrastructure for that to get us an RL codebase that could actually allow us to do large-scale RL and reasoning and this kind of stuff.
Matt Turck [15:22] Great. And then the move to Anthropic.
Sholto Douglas [15:39] And then the move to Anthropic. So the move to Anthropic was in February of this year, and I think it was motivated by a couple of reasons. The number one is I'm just really excited by how deeply every single person in the company cares about how the future goes. And I think that's one thing that really struck me, is everyone at Anthropic has an articulated theory of why what they're working on contributes towards a better future, whether that is AI that is better in ways that can help people improve their lives, or whether it's because it's AI that's safer and more controllable and aligned with our civilization's interests, or even just more deeply understanding what's actually going on inside this AI and trying to better forecast the progress curves and where we think we're actually headed, or policy.
Breaking Into AI Research Without Traditional Credentials
Sholto Douglas [16:17] And probably such a strong advocate for policy in many ways.
Matt Turck [16:41] It's fascinating as a thought, again, seen from the outside of the big AI research labs. A little bit of a question is: how different are those? It seems that everybody's incredibly smart. Everybody has access to the same resources, directionally, more or less. People seem to be focusing on the same problems. And then you see one model comes out and it's better. And then a week later, there's another model that comes out from another lab and is better than the prior one.
Matt Turck [16:54] But from your experience, you see real differences in terms of culture and goals?
Sholto Douglas [17:15] Yeah, I think there are. For example, DeepMind, if you wanted to solve science, is the best place in the world. I think that DeepMind will directly contribute to more scientific discoveries from AI than anything else. Absolutely. And I think it's just so well set up to do this across every aspect. You've got both the direct scientific efforts like AlphaFold and the material science work and this kind of thing, and also generally large efforts to make AI scientists and all that.
Sholto Douglas [17:55] Whereas I think Anthropic has been laser-focused on two things. One is long-term AI alignment, and two is near-term economic impact. So Anthropic has been laser-focused on coding and computer use and things that we think will make a direct impact on the economy within the next six months. One thing that Anthropic noticeably hasn't focused on, compared to DeepMind and OpenAI, is mathematical reasoning. DeepMind and OpenAI have been pursuing mathematical reasoning because of the implications for science and for scientific progress.
Sholto Douglas [18:10] And because I think so many people there just love maths so deeply and would love to see the field progress.
Matt Turck [18:11] Yeah.
What "Taste" Means in AI Research
Sholto Douglas [18:29] We've had to reluctantly sacrifice our focus on that because we want to focus on—well, it's partly for many reasons, but we want to focus on near-term economic impact with the models. And then much of our research along other dimensions is focused on that.
Matt Turck [18:44] Let's double-click on this in a minute. But before we do that, you mentioned a couple of times the word taste, which is one of those important words in 2025. What does taste mean when it comes to AI research?
Sholto Douglas [19:18] I had a really interesting discussion about this with a biology friend where we were comparing taste across biological research and ML. I think one of the most important things is mechanistically understanding exactly what you're trying to do and having an important simplicity regularizer. When you think about taste in ML, it's often the crucial ingredient that allows you to decide what goes into your large training run when you have imperfect information. Because we can study very deeply what the impact of an architectural change is, right?
Sholto Douglas [19:42] But past a point, past a certain level of scale, you have to guess whether or not the impact of that change will compound with other ones, whether it will conflict. Because you can't test your full-scale run, right, N times. You only have one shot at that.
Matt Turck [19:43] Yeah.
Sholto Douglas [19:59] And so a lot of taste comes from being able to make good inferences about, do we think that this ultimately will sort of deliver increasing returns to scale? It also comes down to, do I think this direction of research is worth pursuing? Because often our baselines in ML are so well tuned that it's very hard to beat them, even with what is theoretically a better method, because there are so many small tricks that are required to make a machine learning method work, and they can fail for any number of reasons.
Sholto Douglas [20:41] It's not like building a bridge, where you actually have a pretty good idea why a particular shear was introduced. It can be all these quirks. And so knowing whether it's right to push along that direction or to give it up and try something else is another question of taste. And I think it always comes back to the simplicity regularisation. People love to be clever. We all do. And that's sort of the bitter lesson, I think, is maybe the best expression of this, where generations of people have developed clever methods of encoding priors about how they think an artificial intelligence should reason and encoding it into the model.
Sholto Douglas [21:07] And all of this gets wiped out by scale and through planning, basically like search and learning. And scale applies to those two things.
Matt Turck [21:24] Yes. And the bitter lesson being the Richard Sutton essay, which everybody in AI knows about, but not everyone may have ever heard it, which is exactly what you described: this idea that generalisation and compute will win over time.
Sholto Douglas [21:45] Yes, exactly. The methods that can take advantage of compute, in particular search and learning, will wash away all tweaks. I think I can offer a couple of examples of this in some ways to make it more concrete. One of the reasons that convolutional neural networks were more effective is they encode a prior. And convolutional neural networks, in many ways, you can think of as a little square being drawn across an image, so to speak, that nearby pixels are related to each other.
Sholto Douglas [22:17] And this is a very sensible prior, right? Because if you throw a picture at an AI model and you don't tell it anything about the world, it has to learn that nearby pixels form curves that then form other things. And so there's this hierarchy of abstractions. But ultimately, that is not true of all images. And so convolutional neural networks will be better than a more general vision transformer for the vast majority of images up to a certain amount of scale.
Sholto Douglas [22:58] But past a point, actually, you need to be able to flexibly integrate information across the entire image. And a similar example in language might be, well, we know a lot about grammar, and so you might actually want to decompose a sentence into the constituents—the verbs and the nouns and how they relate to each other and so forth—and provide that explicit structure to your AI algorithm. But then what happens when you want the model to write poetry or to write code? All of a sudden, these assumptions have to be thrown away.
Moving to Google and Building Gemini's Inference Stack
Sholto Douglas [23:06] And so you can't generalize across poetry and code and writing.
Matt Turck [23:26] And to the taste discussion, the art versus science part of this: so are you saying that, at least in terms of anticipating how the training run may go, it's more intuition than actual numbers?
Sholto Douglas [23:46] So you can do actual numbers up to a point. The way to sort of illustrate or think about this is: you are testing a system at multiple levels of scale. And actually, the analogy to biology was, if you think about it, you might test a new therapeutic in a cell and in mice and in model organisms. But that's no guarantee that it will work in a human, right? So you test across multiple different scales and multiple different model organisms, and it seems to work in basic single-cell bacteria, seems to work in mice, maybe works in monkeys.
Sholto Douglas [24:12] That gives you a lot of indication it's going to work in humans, but it's not a guarantee. So at that point, you need to understand the underlying mechanisms of how does this thing work, like what receptors is it binding to, and so forth. And in ML, it's exactly the same, right? You have your different model scales and you figure out, well, okay, it's delivering benefits at these model scales, and I think it should work because mechanistically, I understand what this is doing to the learning dynamics of the model.
Sholto Douglas [24:37] And then you can have confidence that it's going to work. But if it's like, oh, it's a hack and we don't really understand how it works and it's really complicated and introduces all this stuff in the code, then—
Matt Turck [24:47] How often does an idea fail in a company like Anthropic, or in general?
How Anthropic Differs from Other AI Labs
Sholto Douglas [25:09] Yeah, or in general. I mean, I think a good example here is I once asked this question of Noam Chomsky, and he was like, yeah, maybe like 10% of my ideas work. And that's Noam, right? An absolute genius, one of the best in the field. So if only 10% of his ideas work, then I think that establishes a bound on the percentage of ideas that work. Most don't.
Matt Turck [25:31] And it's part of the success, again, of a place like Anthropic or DeepMind to just encourage people to experiment again and again. And I mean, those are very expensive runs, right? Just to say the obvious, a big reason behind the massive amounts of capital going into those companies is that the compute is expensive. And I'm curious about, culturally, the tension between: you need to deliver because there's so much money at stake, versus, no, you should have just a free, open mind and just go for it.
Sholto Douglas [26:09] It's one of the things that at both Anthropic and at DeepMind we really tried to build, which is a culture of safe experimentation where people were trusted to explore ideas for a long time out in the wild, because you often need months of independent research to really prove out a novel research direction that is a substantially different one. It's hard. Particularly, I think it's hard actually less from the compute cost of the experiments and more from a cost of time and focus, because there are so many remaining wins, I guess, even in the current architectures and paradigms and everything. Maybe there's so much low-hanging fruit, right?
Sholto Douglas [26:46] Like, a really high-ROI use of your time would probably be to just go and look at the data and think hard about what the model is learning or doing and make some tweaks to that. You could even just—the simplest things in the world will still deliver massive gains. And so asking people, or giving people the time and space to breathe, and saying, well, we know that there are short-term things you could be doing, but actually we want to try and develop a more general or fundamental technique that allows you to scalably do this in the future, is important.
Sholto Douglas [27:11] There's this tension between doing things at scale and doing things that don't scale, right?
Matt Turck [27:20] Are there people at Anthropic or other places that are deeply researching completely different avenues?
Sholto Douglas [27:48] So, non-transformers, non-RL? Yeah, I think this is another way in which Anthropic and DeepMind differ a little bit. Anthropic is a very focused bet. We think that AGI is within reach in the next couple of years. We think that it's the current paradigms, or something not crazily dissimilar to them. Maybe there's something new, but it's not like we think it's some crazy out-there research program, right? Really, for the last five or six years, Anthropic's ethos has been: scaling compute with broadly the current set of techniques, AGI is tractable within those bounds.
Sholto Douglas [28:16] DeepMind has a much broader scientific culture because it has the resources to do so, right? Anthropic has to be a focused bet. DeepMind has the time and space to be like, well, we're happy to bet on something that is really far outside the current paradigm. And I think, depending on which kind of question you want to ask, whether you think the really focused bet or the wide exploration of different and novel architectures is better, that's one of those sort of research ethos differences.
Sholto Douglas [28:46] Not to say that Gemini itself is a very focused bet, but if you look at Gemini as like 1,000 people, then there's still like 10,000-plus people doing all kinds of really long-term foundational research at DeepMind.
Matt Turck [28:54] Yeah, got it. So, closing the loop on something that you mentioned earlier, why is Anthropic so focused on coding?
Sholto Douglas [29:14] Yeah, we're really focused on coding for two reasons. The first one is that we think it's the thing that will allow us to assist ourselves in AI research fastest. So there's this notion of automating AI research, right? We think that one of the most important signals of the speed of takeoff, the speed of progress, is how much AI is able to assist AI research.
Sholto Douglas [30:05] And so prefetching this is really important, we think. Secondly, we think it's the nearest-term tractable problem domain in terms of economic impact. For Anthropic to be a viable research program that can research the things that we think are important requires economic return. And coding is a huge market full of people who are really, really, really keen early adopters, who love trying and switching things, who are really excited to play with new tools. There's massive, massive demand. There's dramatically more demand for software in the world than there is good software.
Sholto Douglas [30:47] We've seen that in every previous iteration of compilers and general web abstractions and so forth. There's just a booming demand for software. And so, basically, the models are better at coding earlier than anything else because coding is a uniquely tractable problem in some respects for the techniques that we have, in terms of the data exists in many ways. You can containerize and run things in parallel. You can run unit tests, and so you can verify something.
Matt Turck [30:50] You know when it works and you know when it doesn't work.
Sholto Douglas [31:14] Self-driving is uniquely hard, right? You need the car to work first time, kind of. Whereas coding, the models can fail 100 times. As long as it succeeds once, then that's fine. So there's this tractability, there's this replayability that doesn't exist in other fields that touch the real world in some ways. Like, you don't want a lawyer arguing your case that is an AI, right? Because what if it gets the case wrong?
Matt Turck [31:15] Sorry.
Sholto Douglas [31:45] Sorry. Too bad. So, as techniques develop, coding is uniquely tractable. And you can see that, right? Already, I myself am dramatically higher productivity when I'm using the AI tools to write code. And I have a friend who manages nine Claude Code instances, which is just like a crazy number. I don't know how he does that. I can only handle two. So it's maybe like a skill issue on my behalf.
Why Anthropic Is Laser-Focused on Coding
Matt Turck [32:05] All right. Sonnet 4.5 is presented as the best coding agent in the world. So maybe unpack that for us, including performance on the SWE-bench benchmark. What are the numbers? What are the facts? And then we'll go into how that works.
Sholto Douglas [32:32] So SWE-bench is the current benchmark for how we measure coding progress in the outside world, which all the companies use to evaluate against each other. It's an imperfect benchmark in many ways, right? It is like 50% one particular web framework and this kind of thing. But what it does do is it takes real-world scenarios of work that people have done, which is submitting a pull request, so a change to a codebase.
Matt Turck [32:35] And that's stuff that's on GitHub.
Sholto Douglas [33:05] Stuff that's on GitHub, right? And it checks whether or not the model is able to do that same pull request and pass the same tests. And this ends up being a pretty decent proxy for a couple hours of work from a software engineer. These changes aren't incredibly complicated, but they're reasonable complexity, a couple hours of work. We moved recently from roughly 72% to roughly 78% in SWE-bench, which is a pretty substantial step up. I think it's worth pointing out that as recently as a year ago, I think we were under 20% or something like that as a field.
Sholto Douglas [33:43] So there's been dramatic progress in the ability of models to do this unit of work that a software engineer does. I think that SWE-bench is imperfect in a lot of ways, and it's probably pretty close to what we call saturated. One interesting thing to look at with AI benchmarks is you see that they lose their utility past a point because they no longer disambiguate the differences between different models of high capability. But the models, one, are SOTA. They're the best in the world on SWE-bench.
Sholto Douglas [34:04] It's a decent proxy. We're also, I think, more excited by the fact that a lot of our customers and partners are really excited by the model. So one example of this is the Cognition folks and Devin found the model so useful, they had to rebuild their architecture around it.
Matt Turck [34:06] Yeah, they had a great blog post on this.
Sholto Douglas [34:23] Great blog post on this, right? I think that's the real measure of whether or not a model is good, is whether or not it enables people to do things that they couldn't do before. And really, coding as a whole has been transformed in this way over the last year. 3.5 Sonnet, which is the first really strong agentic coding model, the first model that you could ask to do something in front of you, and it sort of was able to interact with your codebase on your computer and do it.
Sholto Douglas [34:48] In many ways, this model is what caused PMF for Cursor. 3.5 Sonnet, because they were in the right place and they were able to capitalize on that model as offering a coding experience that didn't previously exist.
Matt Turck [34:55] Mm-hmm.
Sholto Douglas [35:12] And then actually Cognition and Windsurf went for an even more ambitious target. So basically there's like a spectrum of agency here where either you can ask it to do 30 seconds of work or a couple of minutes of work. 3.5 Sonnet. Then roll into this year.
Matt Turck [35:31] Which, just as a quick aside, is one of the key lessons for anybody in the startup world in 2025, which is: bet on what the models will be able to do in six months from now.
Sholto Douglas [35:59] Right, exactly. Bet on the exponential. So what I think a lot of coding startups are now asking themselves is: what can they now do with models that are capable of independently pursuing goals for substantially longer than previous goals? Before, you had to supervise the models every 30 seconds. Over time, over the next couple of months, you're probably going to end up in a situation where you only need to supervise the models every 10 minutes, 20 minutes or so. That's a pretty dramatic change depending on the complexity of the task.
Sholto Douglas [36:30] Even we have a couple of examples, I think it was mentioned in the blog post, where we asked it to build something that looks roughly like a chat app, something like Slack or Teams, and the model just worked for 30 hours. Like, it was just spinning there on a computer for 30 hours and came out with a really good working Slack-like, Teams-like app. It's pretty incredible. That is nowhere near built into any of the existing products. Maybe Cognition is—Cognition, I think, has always bet on a longer-running, more independent agentic suite.
Inside a 30-Hour Autonomous Coding Session
Sholto Douglas [36:41] And maybe this is the moment that really hits PMF for them, for example. Yeah.
Matt Turck [36:51] Let's unpack the 30-hour aspect, which is fascinating. So first of all, to just ground it for people: so this is computer use?
Sholto Douglas [36:53] Just coding, I think.
Matt Turck [37:00] Just coding. So what does the agent do for 30 hours? Is it clicking on stuff?
Sholto Douglas [37:14] Yeah, it is there. It's reading files and writing code and running tests in exactly the same way that a human would. Basically, you can think of the model as running in a loop where it can constantly decide what to do. People often mention something called tool use, and tool use is the ability to—in this case—use tools like read file, write file, et cetera, or run code in the terminal.
Sholto Douglas [37:55] And it is sitting there in a terminal on a computer in a loop, just constantly looking at the current code, deciding, "Oh, well, it can't quite do this yet, so I'm going to work on that next." It's often making plans, particularly to run for 30 hours. One of the things that we're pretty happy with about the recent launches: we finally taught the models to use what's called memory. And we've built that into the agentic harness. So it's able to create a Markdown file of to-dos and things that it thinks are important to do, check them off, work on them, and check whether they've been completed.
Sholto Douglas [38:36] There's almost this self-verification loop. One of the things that people were worried about with language models over, I think, a year ago or so was that they would fall off track, they wouldn't be able to self-correct, and that this would basically ruin their utility. I think one of the things that's maybe remarkable about the current generation of agents is that they can self-correct. In fact, they're astonishingly good at self-correcting. And this emergent ability has been pretty helpful.
Examples of What AI Can Build in 30 Hours
Matt Turck [38:52] Yep. So much to unpack on this. So this, I think I heard you speak about two axes in the past, one being raw intelligence and the other one being how long an agent can operate.
Sholto Douglas [38:53] Right.
Matt Turck [39:05] So is the fundamental breakthrough, in very simple terms, that if you can do it longer, you basically have a very smart AI that can just work longer?
Sholto Douglas [39:29] Yeah, exactly. If you can maintain long-term coherency, then the model is able to do things that it couldn't possibly have done. If I asked you to just, in a single stream of thought, write a working version of Slack or Microsoft Teams, you wouldn't be able to do it, right? You have to sit there and take notes, do this closed-loop feedback system. So long-term coherency is really important, and it's something that we think is just really critical for this.
Sholto Douglas [40:02] I think a good way of measuring this is to look at the METR evals. They're probably my favourite eval at the moment. And what this eval is, is they've taken a bunch of tasks which humans do, particularly in the machine learning or programming context, and they've annotated how long it takes a human to achieve strong performance at those tasks, and then they ask AI models to do them. And what they found is that there's this really strong relationship between progress and the time horizon over which the AIs are able to complete tasks.
Sholto Douglas [40:28] And so I think it's like every couple of months, the time horizon that the AIs are capable of doing is doubling or something crazy. Maybe every six months the time horizon doubles, which is just utterly insane.
Matt Turck [40:29] Yeah.
Sholto Douglas [40:44] Now, again, like all benchmarks, this one is imperfect, right? It only measures pretty simple tasks. It only measures, I think, 50% success rate or something like this at the task, not like 99% success rate, but it's a good directional measure. And it certainly resonates with my own experiences of, as I've been using the recent models, I start to feel if I just set everything up right, I feel like I could leave this overnight and it could just churn away and it would probably have something pretty useful for me in the morning.
Matt Turck [41:09] Yeah. What are some examples of tasks that you can do with 30 hours that you could not do with shorter runs?
Sholto Douglas [41:40] Yeah, I think in this case, the Slack-like thing is a pretty good example, where it's a significant piece of software. It's really like an end-to-end working piece of software, which often takes a bit of time, like not an MVP demo. Other things I think are interesting is your machine learning experiments and stuff like this are pretty interesting. You want something that's able to propose an experiment and write a bit of code, run some initial tests, come back later, et cetera.
Matt Turck [41:41] Really?
Sholto Douglas [41:49] It opens up the world pretty dramatically. Basically, working software rather than demos, I think, is the key thing.
Matt Turck [41:50] Right.
Sholto Douglas [41:59] Fascinating. Now, I'm not saying that the models will spin you up a full working software right now. It's not going to spin you up a Slack competitor.
Matt Turck [42:04] Although the Claude AI demo that you guys produced was pretty impressive.
Sholto Douglas [42:05] Yes. Right.
Matt Turck [42:25] I think for people who haven't seen it, it shows the progression of the models and how replicating the website went from basically impossible—yes, caricaturing—but like doing wireframes to now doing a fully functional website built autonomously by the AI.
Sholto Douglas [42:54] And it got some pretty complex features. Like, it got Artifacts. Artifacts is a feature where the model is able to write code, and then the results of that code are displayed in the web browser. Claude.ai with Artifacts, with everything else. I can't quite remember how long that one took, maybe a couple hours to do. But basically, regard this as the first halting steps of this. It kind of works. Sometimes it won't work. Sometimes it will. Over the next six months, over the next year, expect dramatic progress here.
Sholto Douglas [43:07] And look at where we are now versus where we were a year ago. And the difference is, I expect the same jump, basically.
The Breakthroughs That Enabled 30-Hour Runs
Matt Turck [43:40] Let's double-click on the breakthrough part of this. Opus 4.1, I think, was able to run up to seven hours. In this case, it's 30 hours, which I realize is not across all tasks, but that's your upper limit. You alluded to some of this memory evolution, this context, this ability to self-correct. Maybe explain in greater detail the advances that enabled that jump to 30 hours?
Sholto Douglas [44:08] I think the biggest things here are, or the question that we often ask ourselves is, what is preventing the models from working for longer, basically? Or when do you need to intervene? And I quite like the model of interventions in a Tesla sense as an example, because right now you need to intervene quite frequently, but it's usually on questions of taste rather than questions of raw programming ability. It's not like the model is unable to, when it's decided to do the right thing, do it.
Sholto Douglas [44:42] But sometimes the models take shortcuts, and sometimes the models forget the overall structure of what they're doing, and they lose themselves in the context. They're doing a locally sensible change, but it doesn't actually make sense in the global context of what they're trying to achieve. And so I think a lot of the improvements, both that we've made and that are still to go, are on this taste and context, basically. It's on making the model better able to decide smart things about the overall structure of the program that it's going to do, and not take shortcuts and write sensible and good code.
Matt Turck [44:55] What about memory?
Sholto Douglas [45:16] Memory is also very important because the models do eventually run out of context. And being able to manage memory over time and even learning from experiences is something which would probably help this a lot. You don't want the model to be constantly rediscovering facts about how a particular system or codebase works. And now this is actually one of those areas where the question of taste or the bitter lesson comes up, because you can imagine us going and launching a massive effort to teach the models coding taste.
Sholto Douglas [45:55] And that could be one way that you solve taste. You have heaps of human software engineers decide, well, no, this is good, or this is bad, or whatever. Where does taste come from in software engineering? Or what do we regard as taste? It's typically that it easily sets you up to make changes later on, or so on and so forth. Or it's easy for maybe multiple agents to communicate with each other and collaborate. Often, good abstractions are something that you or I could work together on in a codebase and not conflict with each other.
What's Actually Driving the Performance Gains
Sholto Douglas [46:28] Right. And so there is this question of how much do you focus on teaching the model coding taste via getting software engineers to decide what is good or bad? Or should you be creating a society of models that all have to together code a giant monolithic codebase, and if they're arguing, then it's bad? You can sort of imagine the spectrum of potential strategies, and picking the right one there is a difficult thing.
Matt Turck [46:36] Sonnet 4.5, again, to the point about the pace of progress accelerating, what were some of the breakthroughs?
Sholto Douglas [47:13] That I can't really talk about. Yeah, I mean, I think it's important to recognize that it's not one individual breakthrough, really. It is the continuous application of lots of different things across the entire stack for many people. And it's mostly just a function of compute in many ways. There are obviously individual breakthroughs, but fundamentally progress has been pretty smooth. Like, on the METR eval, if you look at progress over the last two years, you can plot it with a straight line.
Matt Turck [47:21] Right.
Pre-Training vs. Reinforcement Learning Explained
Sholto Douglas [47:42] And so, similar to Moore's Law of the past and this kind of thing, even Moore's Law is made up of lots of individual improvements. It's not any one critical breakthrough. It's more the accumulation of a huge amount of work in an environment where there's a sort of exogenous force of compute pushing progress forward.
Matt Turck [48:08] Okay, so maybe let's talk about progress at a more abstract level, but grounded in 2025. So a big part of the discussion seems to have been the evolution from a focus on pre-training to RL, which we touched upon a couple of times. Talk about the impact of RL, and why is RL such a big part of the conversation today?
Sholto Douglas [48:35] For those listening, a good way to understand, at a high level, the difference between pre-training and RL: pre-training is like skim-reading every textbook in existence, and RL is like doing the worked problems and getting feedback on whether you were wrong or right. And there are actually a lot of things that you can only learn via RL. A good example of this is the skill to say, "I don't know," in response to a question. Because in pre-training, remember, you're trying to predict what text is going to come next in all of these textbooks, the entire internet in the world.
Sholto Douglas [49:11] And so the only reason you would say, "I don't know," as a pre-trained model is if you think the character that you're modelling in the text would say, "I don't know." Like, if it's a likely completion, right? Not whether you, in fact, don't know, but whether you think that the sort of player that you've pulled from this cast of characters that you could model would say, "I don't know." Whereas in reinforcement learning, you could, in theory, set up a battery of tests where there are things the model knows and things the model doesn't know, and you could reward it for correctly answering things it should know and penalise it for falsely answering when it doesn't know.
Sholto Douglas [49:51] And what it will then learn to do is, it will learn to look up information inside itself and assess its own confidence in whether it knows that information. So saying, "I don't know," or solving hallucinations intrinsically requires reinforcement learning in many ways. So that's one example. There's a whole bunch of things you can't otherwise know. I think also an important change in this sort of era of reasoning models and RL on language models is, at the end of last year, RL on language models finally started to work.
Sholto Douglas [50:26] And I think OpenAI deserves a lot of credit for releasing the first serious RL-plus-LLMs release with o1. And I think this really kicked off a pretty substantial change because it opened up a new axis of scaling. There was pre-training scaling, and now there's test-time compute and RL scaling. I think this is something which all of the research labs were investigating already. One of the reasons that DeepSeek was able to follow so fast was that they'd actually already released papers in the direction of doing RL on language models before, for example.
Sholto Douglas [50:48] And so it was already an idea in the air, but OpenAI deserves the credit for crystallising it, releasing it, and detailing the first public existence of those scaling laws.
Matt Turck [50:57] And maybe to continue making this super educational, how do test-time compute and RL overlap?
Sholto Douglas [51:21] Yeah. One way of thinking about this is test-time compute is doing a lot of reasoning, and then RL is the feedback signal on whether or not that reasoning was right or wrong. And so test-time compute is a way of answering questions that are hard for you to answer. Let's say I ask you a question that you just know off the cuff, off the back of your hand, basically. It's not like from a field that you really know or whatever heuristic that you've already done.
Sholto Douglas [51:46] You've already baked that into your muscle memory, so to speak. But for something which requires you to really think and really learn, like when you're first doing math, if I ask you a basic times table right now, you can say that like this. But if you're a kid, you have to do out the math and all this kind of stuff. You need to do the reasoning chain to learn it, and then you get feedback on whether it's right or wrong.
Test-Time Compute and the New Scaling Paradigm
Sholto Douglas [52:11] So test-time compute lets you do harder problems than you can currently do off the cuff. And RL then allows you to sort of distill that back into the model. It's almost like a ladder. You can constantly do slightly harder problems because you're learning strategies to do harder and harder and harder problems.
Matt Turck [52:35] Reinforcement learning is not a new concept. So we were talking about Richard Sutton, who's been doing work in the field for decades, and others as well. And then there was AlphaGo, that whole line of very successful, impressive RL-based successes. So why is it that in 2025 there seems to be a breakthrough to apply those to LLMs?
Sholto Douglas [52:38] In some ways, it's quite funny.
Matt Turck [52:40] A lot of the—
Sholto Douglas [53:10] Okay, yeah, how do I say it? Let's take the DeepSeek paper, for example. In the DeepSeek paper, they detail, one, an approach that works, and two, a lot of approaches that don't work. Actually, some of the approaches that didn't work were the approaches that led to AlphaGo's success. One of the craziest things about RL on language models in the RL from verifiable rewards regime is it's almost the simplest possible thing. It's almost too simple to work. And this again comes back to that question of taste, where really, I think a lot of people thought this was just too simple to work.
Sholto Douglas [53:44] And so they tried more complex methods that ultimately ended up being harder to get to work. And there may still be juice in those methods, but it was actually really important to nail the simple thing first. And so I think people were almost too ambitious with the RL strategies that they tried initially. I think there's also a minimum bar in LLM quality that is required. You need the model to be able to solve meaningfully difficult coding and math problems before you can get that feedback loop of, well, you solve these ones right and you solve these ones wrong.
Matt Turck [53:58] Right.
Sholto Douglas [54:25] And I think also one of maybe the unintuitive things is those reasoning chains of tokens. People for a long time thought that you'd need to do something clever to give the model long-term coherency. You have to remember that two years ago, 8,000 tokens was a long context for a language model. Five years ago, 8,000 tokens was long. And now models are using 8,000 or 30,000 tokens to reason about something, right? So there was this real phase shift in, oh, language models are smart enough, underlying priors, that they can solve sensibly difficult questions.
Sholto Douglas [54:58] They're actually reasonably coherent at longer contexts than we thought they would be coherent, and this ability to reason in long chains of tokens can emerge naturally with the right feedback signal. And this is a little bit counterintuitive. I think most people wouldn't have expected off the bat that the ability to reason would emerge naturally. There was a lot of thought that you'd have to structure it, you'd have to provide strategies for it to do reasoning, you'd have to build all these things, prompt it and hint it and this kind of stuff.
Sholto Douglas [55:33] It actually turns out, well, no, give it math questions, tell it whether it got them right or wrong, and the model will learn. It comes down to a bit of a lesson in scale and search: just allow the model to search, have enough compute to run the experiments, and the model actually ends up figuring out a really effective and sensible strategy.
Matt Turck [55:39] And that's what's happening now, right? The big labs are basically giving a lot more compute to RL.
Why RL on LLMs Finally Started Working
Sholto Douglas [55:55] Yeah, there's, like, a minimum base model quality, minimum amount of compute for RL, sort of trust in the ability for long-term coherency, doing the simple thing that works. These all sound obvious, but they're actually a little bit counterintuitive sometimes.
Matt Turck [56:08] You mentioned the word AGI earlier. So is your personal sentiment that the combination of ever more powerful LLMs plus RL gets us there?
Sholto Douglas [56:09] Yeah, I think it's sufficient.
Matt Turck [56:16] With the obvious side question of what “there” actually means and what AGI means today.
Sholto Douglas [56:44] Yeah, there's a few definitions that one could use. I think a useful one is better than most humans at most computer-facing tasks, because I think that's a really important moment for the world where we go, okay, intellectual labor is addressable via this set of algorithms, and that totally changes the world. I think there are other definitions that are stronger that you could use. One of those is stronger.
Matt Turck [56:45] That was pretty strong.
Sholto Douglas [57:11] Yeah, sorry. So, I mean, harder to meet, maybe. Yes. Because you could have this and it could still not learn as effectively as humans, right? We learn and generalize from very few examples. We have incredibly high, what's called, sample efficiency, whereas AI models need hundreds or thousands of times more experience, hundreds of thousands of lifetimes, basically, to learn the things that we learn. And they can, over those thousands of lifetimes, learn the skills that we do to an incredibly high degree of accuracy.
Sholto Douglas [57:37] I think one of the important changes over the last year has been that RL has finally meant that we have sort of this algorithm that allows us to take a feedback loop and turn it into a model that is at least as good as the best humans at a given thing in a narrow domain. And you're seeing that with mathematics, and you're seeing that with competitive coding, which are the two domains most amenable to this, where rapidly the models are becoming incredibly competent competition mathematicians and competition coders.
Sholto Douglas [58:13] Right. There's nothing intrinsically different about competitive coding and math. It's just that they're really amenable to RL, more than any other domain. But importantly, they demonstrate there's no intellectual ceiling on the models, right? They're capable of doing really tough reasoning given the right feedback loop. So we think that that same approach generalizes to basically all other domains of human intellectual endeavor, where, given the right feedback loop, these models will get good enough that they are at least as good as the best humans.
Sholto Douglas [58:57] At a given thing. And then once you have something that is at least as good as the best humans at a thing, you can just run 1,000 of them in parallel or 100 times faster, and you have something that's actually, even just with that condition, substantially smarter than any given human. And this is completely throwing aside whether or not it's possible to make something that is smarter than a human. It seems entirely plausible, right? The brain is ultimately a biological computer.
Sholto Douglas [59:29] It seems possible to make a better one. The implications of this are pretty staggering, right? Which is that in the next two or three years, given the right feedback loops, given the right compute, given the right elbow grease and this kind of stuff, we think that we, as the AI industry, are all on track to create something that is at least as capable as most humans on most computer-facing tasks, possibly as good as many of our best scientists in their fields. This is really wild.
Are We on Track to AGI?
Sholto Douglas [59:38] It'll be sharp and spiky. There'll be examples of things it can't do and this kind of stuff, but the world will change.
Matt Turck [59:53] What do you make of the counterthesis of, again, Rich Sutton or Yann LeCun, that seem to be saying that a different approach is needed, or RL only? What do you make of that debate?
Sholto Douglas [1:00:24] Yeah. I think that it's true that our models don't learn anywhere near as efficiently as humans do, right? They take 1,000 lifetimes to learn. But this is, I think, fine because they can live those 1,000 lifetimes, whether in simulations or doing a job at 1,000 firms and so on and so forth. I think that maybe I would disentangle. There's two arguments. One is architecturally that transformers are insufficient. I don't think that's true. I think we haven't yet really found anything that transformers haven't been able to model, provided sufficient data and sufficient compute.
Sholto Douglas [1:00:49] I think RL as an objective is a pretty powerful one. Rich Sutton is actually a big fan of RL as an objective. He just thinks we're actually encoding too many priors in pre-training and this kind of thing.
Matt Turck [1:00:52] It's not an adequate representation of the world.
Sholto Douglas [1:01:06] Yeah, it's not an adequate representation of the world. I think so far the evidence indicates that our current methods haven't yet found a problem domain that is not tractable with sufficient effort. Things that would make me eat my words is if there was some domain that we put a lot of effort into that just didn't move, the benchmarks just didn't move, and we just couldn't make any progress for a year, then I would be like, okay, yeah, there's some fundamental limitation here.
Matt Turck [1:01:29] Yeah.
Sholto Douglas [1:01:55] But instead, what I just constantly see is every time we make a benchmark that measures something we care about, progress is incredibly rapid on that. Yeah, I think this is worth crying from the rooftops a little bit: guys, anything that we can measure seems to be improving really rapidly. Where does that get us in two or three years? I can't say for certain.
Matt Turck [1:01:56] Yeah.
Why the "Plateau" Narrative Is Wrong
Sholto Douglas [1:02:05] But I think it's worth building into respective worldviews that there's a pretty serious chance that we get something that is AGI.
Matt Turck [1:02:26] So you think people don't realize—it's always interesting, right? Because reading stuff online in the last three, four months, it's like this theme of we've reached a plateau. But you're basically saying the opposite, right? We are in an exponential curve, and many people don't realize that it's the case.
Sholto Douglas [1:02:50] Exactly. And I mean, people have said that we're hitting a plateau every month for the last three years. And if you look at where we've come over the last three years, it's incredible. I think that one other thing that makes me think, God, we're not anywhere close to a plateau, is I look at how these models are produced, and every part of it could be improved so much. It is a primitive pipeline held together by duct tape and the best efforts and elbow grease and late nights.
Sholto Douglas [1:03:19] And, God, actually, I remember—I don't know if this is a good analogy or whatever—but I remember I went sailing with a couple of friends a few months ago. And the boat was so well designed. It was just clearly the product of millennia, or centuries, of accumulated human design and effort. And I was like, wow, this is what it feels like to be in the accumulation of a lot of human effort, right?
Sholto Douglas [1:03:38] It's actually pretty hard to beat today's best sailboat designs. Five years of best effort, last-minute desperate effort, and there's just so much room to grow on every part of it.
Sonnet's Performance Across Economic Sectors
Matt Turck [1:03:57] Sonnet 4.5, which is described as the best coding model in the world, also seems to be performing across a lot of different other domains, like economics, research, and finance. So it's early, just to verbalize that.
Sholto Douglas [1:04:17] One of the things that I was really excited by, actually, was that GDPval that OpenAI released. Claude Opus 4.1 was the leading model there. And I think that's a really interesting and really good eval because it demonstrates such a breadth of tasks across all parts of the economy.
Matt Turck [1:04:40] It's an eval that's across the various sectors of the economy, right? So manufacturing. And basically, they took a bunch of experts to describe what success looks like. And now the models are going to be able to be measured not across just coding or some limited tasks, but across everything. Is that a fair way of describing it?
Sholto Douglas [1:04:59] I've wanted someone to do this for a long time. Take the Bureau of Labor Statistics. I think the most important input into policy would be: take the Bureau of Labor Statistics, take all the jobs there, break them down to tasks, and see whether the AI models are able to do that, and measure progress over time. Right. And this is obviously going to be an imperfect measure. We'll probably reach better than human on GDPval, and it won't change anything economically because it'll be all the connective tissue and all the context, and actually the tasks won't be representative.
Sholto Douglas [1:05:26] But again, we'll then find better ways to measure these difficulties, and we'll keep pushing benchmarks. I've wanted someone to do this for a long time. I'm really glad they did it. I'm really glad that our models were general and generally strong and sort of showed up top across all the areas. And I think policymakers should really look at this and extend it and really make an effort in investing in figuring out whether we are on track for what I've been claiming we're on track for.
Sholto Douglas [1:05:38] Right. We can measure this.
Matt Turck [1:05:38] Yes.
Sholto Douglas [1:05:39] And we should be.
Preparing for a World of 10–100x Individual Leverage
Matt Turck [1:05:55] Yes. So, yeah, just to close on that last theme: so awesome, all very exciting. What do we all do? How do we prepare for this world that seems to be around the corner?
Sholto Douglas [1:06:19] Yeah, I think the most actionable piece of advice is: keep planning for a world where you, as an individual, have more leverage. Right now, I can use two coding agents to do twice the work that I could have done before. If coding agents progress in the way I've been saying, in a year or two, you'll be able to manage a team, basically, that works 24/7 for you, doing work. I think we should expect, in the digital domain, for individuals to get dramatically more leverage over the next couple of years.
Sholto Douglas [1:07:00] I think there are many incredibly important problems we're going to try to—our world is so imperfect in so many ways. People still live in dramatic poverty. Health and medicine are unsolved. Housing is completely unsolved. The world could be a million times better in so many different ways. And what I hope is that people take, initially, models giving us leverage over the digital world, and then hopefully models giving us leverage over the physical one through robotics, to dramatically improve it.
Matt Turck [1:07:14] Is that happening? Robotics? That's another thing that seems to be one of the key themes. But on the other hand, to use actually the word hand, it seems like people are still struggling to make hands move.
Sholto Douglas [1:07:15] Right.
Matt Turck [1:07:18] The way—so the physics of it seems to be the limiting factor. Yeah.
Sholto Douglas [1:07:41] There's this thing called Moravec's paradox, right? Which is that things which we find really easy, like manipulation, picking up objects, are really hard for AI. But maybe things which we find hard, like reasoning through mathematical problems, are easy. I actually think Moravec's paradox is a little bit fake, and I think this is mostly a question of data availability and RL signal and stuff. And I think one interesting reason to look at this: if you look at robotic locomotion, so the ability of robots to walk around and balance and stuff, and look at the videos of the Unitree robots, the difference now versus two years ago is crazy.
Sholto Douglas [1:08:19] These things are incredibly agile. There's this video, I think, of someone kicking one over, and it literally does a Matrix kind of get-back-up thing. It's crazy. This is because locomotion is a really easy RL signal. And right now, locomotion is kind of solved, to be honest, with basic RL. Manipulation is a bit harder, but there's a few things that make me think that robotics is going to work. For starters, I've seen incredible progress from the robotics labs over this year.
Sholto Douglas [1:08:47] Really, they've gotten to the point where they can do pretty interesting basic physical tasks. Two is the existence of a large generator-verifier gap, which is that one of the things that makes improving our models hard is we constantly need to find people who can beat the models at the things we want to improve them on. But with robotics, we're making really smart general models. So you can actually have these as teachers or judges for whether or not the robot is doing the right thing.
Sholto Douglas [1:09:19] If I say, "Stack the red block on top of the blue block," we can then ask the language model, "Did it stack the blocks appropriately?" If so, give a reward. If not, don't. So you can use the generator-verifier gap to give models feedback. And finally, for a long time in robotics, people thought they would have to solve long-term coherency and planning. And that's also something that language models have made easier. They can break things down into multiple steps. So all the robotics labs are focused really hard on making great motor policies, and they're making incredible progress.
Sholto Douglas [1:09:30] It's mostly just a data and feedback loop question.
Matt Turck [1:09:42] All right, Sholto, it's been fascinating. I can think of another 40 questions that I would want to ask you right now, but you've been incredibly generous with your time. Thank you so much. This was terrific. Really appreciate it. It was a real pleasure.
Sholto Douglas [1:09:42] Thank you very much.
Matt Turck [1:10:03] Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.