Are We Misreading the AI Exponential? Julian Schrittwieser on Move 37 & Scaling RL (Anthropic)

The MAD Podcast with Matt Turck · with Julian Schrittwieser, AI researcher, Anthropic

Julian Schrittwieser is the AI researcher at Anthropic. We cover why benchmark trends suggest top models can independently complete a full day of work within two years, why RL makes agents robust by training on their own failures and recoveries, and why he expects models to reach Nobel-level scientific insight by 2027 or 2028.

Watch on YouTube

Chapters

  1. 1:09 — The “exponential” from inside frontier labs
  2. 4:46 — 2026–2027: agents that work a full day; expert-level breadth
  3. 8:58 — Benchmarks vs reality: long-horizon work, GDP-Val, user value
  4. 10:26 — Move 37 — what actually happened and why it mattered
  5. 13:55 — Novel science: AlphaCode/AlphaTensor → when does AI earn a Nobel?
  6. 16:25 — Discontinuity vs smooth progress (and warning signs)
  7. 19:08 — Does pre-training + RL get us there? (AGI debates aside)
  8. 20:55 — Sutton’s “RL from scratch”? Julian’s take
  9. 23:03 — Julian’s path: Google → DeepMind → Anthropic
  10. 26:45 — AlphaGo (learn + search) in plain English
  11. 30:16 — AlphaGo Zero (no human data)
  12. 31:00 — AlphaZero (one algorithm: Go, chess, shogi)
  13. 31:46 — MuZero (planning with a learned world model)
  14. 33:23 — Lessons for today’s agents: search + learning at scale
  15. 34:57 — Do LLMs already have implicit world models?
  16. 39:02 — Why RL on LLMs took time (stability, feedback loops)
  17. 41:43 — Compute & scaling for RL — what we see so far
  18. 42:35 — Rewards frontier: human prefs, rubrics, RLVR, process rewards
  19. 44:36 — RL training data & the “flywheel” (and why quality matters)
  20. 48:02 — RL & Agents 101 — why RL unlocks robustness
  21. 50:51 — Should builders use RL-as-a-service? Or just tools + prompts?
  22. 52:18 — What’s missing for dependable agents (capability vs engineering)
  23. 53:51 — Evals & Goodhart — internal vs external benchmarks
  24. 57:35 — Mechanistic interpretability & “Golden Gate Claude”
  25. 1:00:03 — Safety & alignment at Anthropic — how it shows up in practice
  26. 1:03:48 — Jobs: human–AI complementarity (comparative advantage)
  27. 1:06:33 — Inequality, policy, and the case for 10× productivity → abundance
  28. 1:09:24 — Closing thoughts

Transcript

The “exponential” from inside frontier labs

Matt Turck [1:07] Hey Julian, welcome.

Julian Schrittwieser [1:09] Hey Matt, thanks for having me.

Matt Turck [1:25] A couple of weeks ago, you wrote an incredible blog post that broke the internet, entitled “Failing to Understand the Exponential, Again.” What is it that so many people are missing about the current trajectory of AI?

Julian Schrittwieser [1:42] Yeah, it's funny that you bring up that blog post. I really didn't expect it to blow up that much. I actually had the idea when I was on holiday in Kyrgyzstan a few weeks ago, on a very long car ride. And then I started thinking about this and all the talk about AI bubbles I had seen on X, and this discussion. And it seemed very divorced from what was happening in frontier labs and what we were seeing.

Julian Schrittwieser [2:05] And that made me start to wonder a bit: is it that things are moving so fast that people maybe struggle a bit to extrapolate and understand intuitively? Maybe it's far away now, but it's doubling every so many months, which means that once it gets close to us, it's going to move past and become really good very quickly. And that reminded me a lot, in a different way, of what happened during early COVID, where we had a similar situation where at the beginning it's very few cases.

Julian Schrittwieser [2:24] It's like, well, it's never going to happen. It's only a few hundred people. Who cares? But if you understand the math and if you look at it right, it's like, oh, it's going to double every week, two weeks.

Matt Turck [2:24] Yeah.

Julian Schrittwieser [2:40] Clearly it's going to be a massive scale, but it's very hard for us to intuitively understand these exponential trends because it's just not what we're used to in our normal environment. And so that's what got me thinking: is something similar happening here with AI, right?

Matt Turck [2:41] Yep.

Julian Schrittwieser [3:00] We are clearly, if you're looking at many benchmarks we have, many evaluations we have, seeing this very consistent improvement over many, many years, where every, say, three, four months, it's able to do a task that is twice as long as before, completely on its own. And so we can extrapolate this, right? And we see that in a year from now, maybe two years from now, the top models are going to be able to work completely on their own for a whole day or more.

Julian Schrittwieser [3:35] Combined with the fact that there's a huge number of knowledge-based jobs in the economy, knowledge-based tasks, and combined with the fact that in the frontier labs, we are not seeing any slowdown of progress, just extrapolating those things together over a very short time, like half a year, one year, already that is enough to know that there is going to be massive economic impact. That means if you look at OpenAI, if you look at Anthropic, if you look at Google, those valuations, those revenue numbers are actually fairly conservative.

Julian Schrittwieser [4:12] I think some more thoughts, some things I've seen more recently, is that it's maybe actually even more interesting and more complex: while those frontier labs and frontier models are clearly very capable and on an extreme trajectory, there are a lot of other companies that are trying to follow into the same AI sphere that may also have very high valuations, but not necessarily the revenues to support it. And so it's possible that there may simultaneously be some sort of bubble in the wider ecosystem while, at the same time, the frontier labs are on a very solid trajectory, having a lot of revenue, making a lot of money.

2026–2027: agents that work a full day; expert-level breadth

Julian Schrittwieser [4:46] I think that may be quite an unusual situation that in the past, maybe in the dot-com bubble, maybe people were talking about the railroad rush and stuff like this, we did not see this bifurcation. So I think, yeah, I've been thinking about this more, and I think the situation is getting more and more interesting.

Matt Turck [4:56] Fascinating. So you alluded to some of your predictions or extrapolations for '26 and '27. Do you want to unpack that? You had three of those.

Julian Schrittwieser [5:21] Maybe calling it my prediction is giving myself too much credit. I will just say this: if you look, for example, at METR eval and you very naively extrapolate the linear fit, that's what you would expect to happen. And so I'm just going to be humble and say, most of the time, I'm not going to be smarter than statistical models, statistical extrapolation of past trends that have been very consistent. So I'm just going to be very humble. Despite what I might know about research and what's happening, probably the most likely, the best prediction I can make is actually just follow that data, that extrapolation, and see where it's going to take us.

Julian Schrittwieser [5:57] And yeah, in that case, if you roll this out, if you look at other benchmarks, I think we would have something like next year, maybe the models will be able to work on their own for a whole day's worth of tasks. If you think of software, you might say, like, implement this entire feature, build out this entire set of the app. If you think of knowledge work, maybe do a whole research report, this kind of scale. The reason I think task length specifically is interesting is because that's what allows you to delegate more and more work to language models, to agents.

Julian Schrittwieser [6:28] Even if you have a very clever model, but if it needs feedback or interaction with you very often, then it really limits what you can delegate to it. If you need to talk to it every 10 minutes, right? Versus if you have something that can go for hours at a time, obviously, right? Then you cannot just have one copy of it. You can have a whole team that you delegate tasks to and manage. And so I think that's why it's really critical that the models are actually smart enough, the agents are smart enough, to work on their own, to correct their own errors, to iterate, because that's really what allows you to delegate.

Matt Turck [7:16] Indeed. Task length and time to complete as the metric for progress. So by mid-2026, you mentioned agents can work all day autonomously. Late 2026, at least one model matches industry experts across many occupations. And then by 2027, models frequently outperform experts on many tasks. So it's more time running and then generalization across the economy. And you mentioned GDPval, the OpenAI metric, as a benchmark to already see the progress towards multiple professions.

Julian Schrittwieser [7:45] Yeah, I think GDPval is a super cool evaluation from OpenAI, where they collected a lot of real-world tasks from real domain experts to make sure it is actually representative of what you might do in the economy. And then they evaluated a lot of models on those tasks. They compared them against real experts' performance to give us a really good indication of how close, how far are we from having significant economic impact. So I think that's a super cool evaluation.

Matt Turck [8:00] The sort of obvious question is that GDPval and METR are carefully designed benchmarks. How do they predict production value once you add compliance, liability, messy data, the messy world, tool friction, and all the things?

Julian Schrittwieser [8:28] So I think messiness and task length, time duration that you're able to work independently, are very similar or very correlated. So I think that's why it's interesting that METR tries to measure how long can the model go on its own. Because if you think about how do you come up with a task that takes a human eight hours, 16 hours, you will have to include all this messiness and all this real-worldness to even be able to measure it. But I think ultimately, to go further, we really want benchmarks.

Julian Schrittwieser [8:57] We really want evaluations that come from the actual users, whether it's the industry, whether it's private users, because that's what ultimately matters. Is the model helpful to you? Do you get something out of it? Does it do your paperwork, help you write something, fix your code, help you study? I think that's the real proof. If you release a new model, do people start using it more?

Benchmarks vs reality: long-horizon work, GDP-Val, user value

Matt Turck [9:15] Is there anything that would change your mind? Any kind of signal, whether that's real-world adoption or benchmark performance, something that would make you more cautious about that exponential? Is there anything that would change your mind?

Julian Schrittwieser [9:37] I mean, many things, yes. I think many of these things are internal only. I might look at our model pre-training. I might look at our fine-tuning. I might look at the RL thing. How do new runs go compared to past runs? Do they match our expectations? Does scaling continue? Then I might look at more public things, like are people actually able to use those models to be more productive, for example? At the beginning, there's always some adaptation period of, oh, you have a new tool like Claude Code.

Julian Schrittwieser [9:57] It takes you some time to figure out how to use it. But then in the medium term, in the long term, do people keep using it? Are they getting more and more productive using it? I think that's one of the things I look at. Many, many signals. I think when you do RL, when you do research, you get very much in the habit of looking for signals to prove yourself wrong because you often have ideas that you get attached to, but that's not a good way to do research.

Move 37 — what actually happened and why it mattered

Julian Schrittwieser [10:27] Most of your ideas are not good, and they're not going to work. So you really want to figure out as quickly as possible whether this idea is any good or whether it's actually wrong. So you really get into this habit of finding the fastest thing that will show that, oh no, this is actually not true.

Matt Turck [10:54] So by 2026, 2027, in your extrapolation framework, AI becomes as good as humans. A key question of the moment is: to which extent can it become better than humans? There's some chatter these days around Move 37 and whether AI can create those alien new paths to think and solve hard problems. So first of all, maybe remind the audience what Move 37 is, and then do you think that AI in its current state is going to be increasingly able to provide Move 37-type thinking?

Julian Schrittwieser [11:31] To give background, Move 37, that was when we were building AlphaGo, an AI program to play the game of Go. So it was in the year 2016, I think, and we were playing one of the best players in the world at the time. Because at that time, no AI program, no computer program had ever beaten the top human players at Go. And it was considered to be one of the most difficult board games, sort of like a real test of intelligence.

Julian Schrittwieser [11:57] Move 37 happened during the second game of the five-game match, where AlphaGo played a really unexpected, unconventional move that surprised many professional Go players. I think the commentator said it was truly creative, unexpected. And then ultimately AlphaGo ended up winning that game. And so I think that was, for many people, an early sign that AI is not just purely calculating, following an optimal path, but it can also do something that is truly novel and creative that you might not expect just from imitating its training data.

Julian Schrittwieser [12:33] I think that's very relevant in the modern context as well, right? Because, as you alluded to, there is a lot of discussion of, oh, are LLMs just parroting the training data? Can they actually do novel things? For me, as somebody who has been doing research a long time, I think it's pretty clear that these models can do novel things. And that's why they're so useful to many people, whether it's writing code for you, because obviously, you're not just writing code that you already have.

Julian Schrittwieser [13:09] That wouldn't be very interesting. Or helping you write a paper. The way those models are trained, they're literally trained to generate a whole probability distribution, which means that when we sample from them, we can generate an infinite number of novel sequences from them. For the question of something like Move 37, I think there it really comes down to: is it something that is sufficiently creative and impressive?

Matt Turck [13:09] Yeah.

Julian Schrittwieser [13:35] That we can easily recognize it. In the game of Go, that was pretty ideal conditions because it's very clean, very abstract. Each move is very impactful, so you can really see it clearly. I think to have the equivalent for our modern models, you need the combination of a task that is sufficiently difficult and interesting and a model that is both able to create sufficiently diverse and creative ideas, and also able to evaluate accurately how good they are so that it can go down increasingly novel paths while making sure that this novel path is actually interesting and useful.

Novel science: AlphaCode/AlphaTensor → when does AI earn a Nobel?

Julian Schrittwieser [13:55] Creating novel things is actually very easy with language models. The hard part is creating novel things that are useful and interesting.

Matt Turck [14:27] Extrapolating this further, there's the idea of creating novel science. So not just one move, but a whole new idea, new concept. Current take on this? So I think AlphaCode and AlphaTensor proved that you can discover novel programs and algorithms. Very recently, I think last week, there was some news of Google DeepMind and Yale in the biomedical field coming up with brand new things as well. So do you think that's accelerating, and that AI is in the process of discovering novel science?

Julian Schrittwieser [15:05] I think we are absolutely at the stage where it is discovering novel things, and we're just moving up the scale of how impressive, how interesting are the things that it is able to discover on its own. I think it's highly likely that sometime next year, we are going to have some discoveries that people pretty unanimously agree are super impressive. I think at the moment, we're more at a stage of, oh, it came up with something, but there's debate about it. But yeah, I'm not very worried, because I see this process continuing, and then once it gets clear enough, there's less need to argue about it.

Matt Turck [15:18] How far do you think we are from an AI winning the Nobel Prize?

Julian Schrittwieser [15:41] I think that's a really interesting question because we had a Nobel Prize for AI with AlphaFold, of course. And so I think the next very interesting point is going to be: when can AI, on its own, make a breakthrough that is so interesting that it would win a Nobel Prize? I think my guess for that level of capability might be maybe 2027. I think we're probably not going to find out for quite some time afterwards because of the delay in getting prizes.

Julian Schrittwieser [16:01] But I think by 2027, 2028, it's extremely likely that the models will be smart enough and capable enough to actually have that level of insight and that level of discovery.

Matt Turck [16:02] Amazing.

Discontinuity vs smooth progress (and warning signs)

Julian Schrittwieser [16:26] I think not just Nobel Prize, right? It's like the Fields Medal for math and all these kinds of advances. I think that's what I'm truly excited about, actually, is AI that can help us advance science and really unlock both all the mysteries of the universe and all the improvements in living standards and abilities for us that we could have if we understood the world better.

Matt Turck [16:47] So extrapolating this even further, then we get into the AI 2027 thing that you probably saw. So this general idea of if AI can create novel science, then AI can create AI researchers, and basically AI can iterate itself, which effectively leads to a discontinuity moment. So I don't know if that in the blog post, that's the singularity or whatever, but does that strike you, as somebody who's as deep in the field as possible, as something that is possible in the short term, or are there counterbalancing forces that make that path to discontinuity harder as you get closer?

Julian Schrittwieser [17:41] Yeah, I think a true discontinuity is extremely unlikely. Obviously, AI researchers are already using AI to accelerate themselves. And so what's already happening, and what is likely to continue happening, is that we see a smooth improvement of productivity. And then the main open question is: how does the difficulty of improving AI keep scaling? Because a very common effect and a very common issue in many scientific fields is that we find all the easy problems first. And then, as we continue exploring the field, it gets more and more difficult to make advances.

Julian Schrittwieser [18:17] So in my mind, the main question is: do these two trends balance each other out so that the AI makes us increasingly more productive, so that as it gets more difficult to make advances, we just about stay on trend and then we keep improving roughly linearly? Or is it still too difficult, and then eventually, after some time, we still see a slowdown? But it seems quite unlikely to me that we improve in productivity so much that we can actually accelerate. That would be very unlike any other scientific field.

Julian Schrittwieser [18:46] The normal course in many scientific fields is that we actually need to exponentially increase the research effort just to keep making progress and find new insights. For example, if you look at pharmacology, discovering new drugs, it's nowadays in the range of billions of dollars to discover a new drug, versus maybe 100 years ago, a single scientist could discover the first antibiotic by accident. It's not that we will be surprised by a sudden takeoff in progress where we're just doing our research and suddenly our model is 10x better.

Julian Schrittwieser [19:07] We will be seeing advance signs of, oh, we are making faster progress every single week. We can see something is happening. Maybe we decide to pause if we don't understand what's happening.

Does pre-training + RL get us there? (AGI debates aside)

Matt Turck [19:37] Do you think the current approach to modern AI systems, which effectively is pre-training plus RL, does that take us to where we want to be? Whether we call it AGI, ASI, it's unclear what any of those things mean. But do you feel that this paradigm is the right one, or do we need to come up with a different architecture altogether, post-transformers or otherwise?

Julian Schrittwieser [20:00] I think that's a great question. And I think it hugely depends on what you mean by where do we want to be. So I think if you're thinking, oh, we want some kind of system that can perform at roughly human level in basically all tasks that we care about productivity-wise, then I think, yeah, it's extremely likely that the current approach—pre-training, RL, transformers—is going to get us there. If what you care about is, oh, we want to have a model of intelligence that is conscious in the same way we are, or more abstract qualities like this, I think that's maybe more uncertain.

Julian Schrittwieser [20:28] And I think this is where a lot of the confusion and disagreement comes from. As you alluded to, AGI, ASI—people talk about very different things, and they have very different things in mind when they say, oh, the current paradigm is going to get there; it's not going to get there. I often like to not use the term AGI or ASI and just talk very concretely about what problem are we solving, what task are we solving, what quality are we interested in, because I find that often makes the actual disagreement much more obvious.

Sutton’s “RL from scratch”? Julian’s take

Julian Schrittwieser [20:55] But yeah, I think if you're just thinking in terms of, is this going to help us be massively more productive? Is this going to massively accelerate scientific progress? Then I think definitely the current approach will get there.

Matt Turck [21:25] And given how extremely deep you are in RL, I cannot resist asking you the trendy question du jour, based on Richard Sutton's recent appearance on Dwarkesh's podcast. Do you think that the models of the future will be trained in RL from scratch, and that actually having pre-training in addition to RL is the wrong way to go?

Julian Schrittwieser [21:55] Personally, I think that's unlikely. Not because pre-training is strictly necessary. I think we may well be able to train something completely from scratch, as we've been able to do in other domains, but more because pre-training on these vast datasets that we have just brings us so much value that we would, from a practical point of view, not want to give it up. So we might well do some agents that are trained from scratch out of scientific interest. It could be very interesting to learn about what a non-human intelligence would look like.

Julian Schrittwieser [22:26] But from a pragmatic point of view, I definitely think we would keep using pre-training data, not just from an efficiency point of view as well, but also because I think there are interesting safety angles. Because by pre-training on all this human knowledge, we're implicitly creating an agent that has similar values as we do. And I think that is quite valuable for aligning a highly intelligent agent. If you already start out by caring about the same rough set of values, that makes things much easier than if you create an arbitrary alien intelligence that may have completely different values.

Julian Schrittwieser [22:46] Despite having done a bunch of from-scratch RL in the past, I think I'm often quite pragmatic about this.

Julian’s path: Google → DeepMind → Anthropic

Matt Turck [23:20] I'd love to put a pin in that specific discussion about alignment and what we do to ensure safety until later in the conversation, because I think that's a super interesting vein. But maybe to switch tacks for a minute, I'd love to go into a little bit of your story and then the monumental body of work that you've done at Google DeepMind before joining Anthropic around AlphaGo, AlphaZero, MuZero. So maybe just the three- to four-minute version of your personal story from when you were a kid.

Matt Turck [23:34] What was the path that led you to become a world-class AI researcher?

Julian Schrittwieser [24:01] Yeah, actually, when I was a kid, I didn't have any expectations of becoming an AI researcher. I was always very interested in computers, and I grew up in the Austrian countryside in a small village. So it's not like there was a huge amount of things happening, but computers were always very interesting to me. It's like this connection to the wider world, to all these other interesting things. And I was very interested in computer games as well. And I think that's the first time I became interested in programming because I wanted to make my own games.

Julian Schrittwieser [24:32] Which I think is very common in people who get into programming. But I somehow always got distracted by the technical aspect of, oh, I'm gonna build a very general game engine that can run any kind of game. And so I never actually ended up making any game. I learned a lot about making game engines and different technologies. And that's how I ended up studying computer science eventually in Vienna. Yeah, that was like a classical computer science degree.

Julian Schrittwieser [25:01] And then by chance, after my first year, in my first summer holidays, I had an internship with Google. And that's when I realized, oh wow, these guys are doing really interesting things. That's where their big clusters, the tens of thousands of machines, are. That's the first time I radically changed my plans from wanting to stay in academia, and I had originally thought, oh, maybe I'll do a PhD. That's when I changed. Oh no, actually, I just want to join these guys at Google, and I will finish my degree as quickly as possible.

Julian Schrittwieser [25:36] And so that's when I got my full-time position at Google, finished my degree the next year, and then moved to London. So I was just working as a normal software engineer at Google, working actually in advertising, which I wasn't super excited or interested in. So the technology was interesting, right? It's like these huge systems, and Google has famously great technology. But actually, after a year-ish of this, I was pretty done and bored of advertising. And so I was actually planning to leave Google and thinking of maybe joining a hedge fund, going into finance, when by chance I saw an email in my work inbox that this guy, Demis, was going to come to the office and give some talk about—oh, Demis Hassabis.

Julian Schrittwieser [26:05] Atari and video games and AI. And it was actually a day off because I was visiting a friend somewhere else in England. But that email looked so intriguing that I was like, oh no, I'm going to have to take the train back to the office right now.

Matt Turck [26:06] Amazing.

Julian Schrittwieser [26:25] And see this talk. And yeah, I'm really glad that I saw this email and I did go back, because that's the moment where I decided, oh no, I'm not going to go into finance. I'm going to move to DeepMind. I'm going to join these guys because this looks clearly super interesting, super amazing. They are doing really interesting research.

AlphaGo (learn + search) in plain English

Matt Turck [26:56] All right, tell us the story of AlphaGo, AlphaGo Zero, AlphaZero, MuZero—what those are—because it feels like it's fundamental AI knowledge that everybody who has an interest in the space should know about, should understand the progression in particular. So, starting with the beginning of AlphaGo, you alluded to it a second ago, but what did it do? How was it trained? And then how did that evolve with each version?

Julian Schrittwieser [27:24] AlphaGo, I think at that moment in time, Go in the machine learning community was this really big target where everybody felt like, oh, it's this big unsolved challenge. ImageNet had just happened before. So clearly deep models were starting to do something with images and being able to recognize them and predict them. And if you look at the Go board the right way, it looks a lot like one of those images that you classify. So there was a lot of momentum around using neural networks to somehow play Go.

Julian Schrittwieser [27:48] And then, at the time, David Silver and Aja Huang at DeepMind had been working on Go. I think both of them had been working on Go for quite a while, had published some very interesting papers. And that's when the idea of using Monte Carlo tree search with deep networks came together. So the idea was to train a deep neural network to predict which moves you might want to play and whether you're winning or losing the game, and then use the tree search to really make a big plan of what are all the possibilities in the game.

Julian Schrittwieser [28:12] How would it go for you if you chose a certain move or a different move? How would the opponent respond?

Matt Turck [28:27] And to explain this in super plain English, the term search in this case is, as you said, tree search. It's not what people normally think of as search, which is searching a corpus. This is searching a series of options, effectively. Is that the right way to think about it?

Julian Schrittwieser [28:45] Yes. It's quite literally what you might do when you play a game of chess, when you play any board game. It's quite literally thinking of what move am I going to do, what move is my opponent going to do in return, and then thinking about many possible moves like that and mapping out all the possibilities in the future.

Matt Turck [28:49] So, deep learning plus search. What was AlphaGo trained on?

Julian Schrittwieser [29:18] The first training phases of AlphaGo were on some human amateur games, if I remember correctly. So basically, if you have humans playing many games of Go, try to predict at each turn in the game what move they would have played. And it turns out that if you train a deep network to do that, you can get something pretty decent, like amateur Go level, but not good enough to actually beat a really strong player.

Matt Turck [29:32] And by the way, just for the lore of it, did you guys have any sense that AlphaGo was going to crush Lee Sedol, the famous Go player that you mentioned earlier in the conversation? Was it obvious before, or was it a surprise?

Julian Schrittwieser [29:51] We thought we had a pretty good chance, but we were very nervous about, like, are we going to win or are we not going to win? Are we going to lose? Yeah, we actually had some bets beforehand of, like, how many games are we going to win or lose? I think it was very ambitious to put the match as early as we did. If we had wanted to be a bit more safe, we may have tried to do it a few months later.

Julian Schrittwieser [30:06] And I think if we had done it a few months earlier, we would have probably lost. So it was very knife-edge, I guess, which also made it much more interesting for us, right?

Matt Turck [30:07] Yeah.

AlphaGo Zero (no human data)

Julian Schrittwieser [30:17] Because it really means that each game is like a nail-biter of, oh, what's going to happen? Are we going to win? Are we going to play a dumb move? What's going to happen? So that was very exciting.

Matt Turck [30:23] AlphaGo Zero, which was, I believe, the year after, how was that different? What was the progression?

Julian Schrittwieser [30:45] The main change between AlphaGo and AlphaGo Zero was to remove all the human Go knowledge. So instead of starting by imitating human Go games, we were training it just from scratch, playing only against itself and rediscovering basically all Go knowledge, completely figuring out from scratch how to play.

Matt Turck [30:47] Did you give it the rules of the game?

AlphaZero (one algorithm: Go, chess, shogi)

Julian Schrittwieser [31:00] We didn't give the rules of the game to the network per se, but we used the rules of the game to score the result. So basically, it would play, and then we would tell it who won, who lost, or, you know, you cannot make this move.

Matt Turck [31:07] So the next hop was AlphaZero, which was a year or two later.

Julian Schrittwieser [31:07] Yes.

Matt Turck [31:08] How is that different?

Julian Schrittwieser [31:35] So AlphaZero, the idea was, well, obviously Go is a really beautiful game, but ultimately we would like to do something more general, right? So can we remove anything that is Go-specific and verify that the algorithm can actually solve more problems? And in that case, we did that by trying to solve chess, Go, and shogi, which is basically Japanese chess, with the same algorithm, same network structure.

Matt Turck [31:36] Yeah.

MuZero (planning with a learned world model)

Julian Schrittwieser [31:47] Just by running it in different games and also making it much simpler, elegant, faster. So basically, that was really laying the groundwork for applying the algorithms to solve real problems.

Matt Turck [32:12] And then the next stop in the journey was MuZero. And just to bring it home for people, you were, I believe, second author on AlphaGo Zero, and you were the lead author on MuZero, which, in the world of AI—I'm sure you're going to be very humble about it—but in the world of AI, it's as big a deal as it gets. So I'll say it so that you don't have to say it. So, MuZero, what was the next—how was that different?

Julian Schrittwieser [32:36] So the main motivation I had for making MuZero was that if you want to solve many real-world tasks, you have no way of perfectly simulating what's going to happen. If you play a board game, obviously, if you make this move, you know what's going to happen. It's like the piece is going to go there, it's going to take a piece, whatever, right? But if you actually want to solve something like a robotics task or anything more complicated, it's impossible for you to simulate what's going to happen accurately.

Julian Schrittwieser [33:12] And also, we as humans, we don't do this, right? We just imagine in our head, oh, if I'm going to say this, then he's probably going to respond in that way. This meant that AlphaZero, as it was, could not be applied to such problems because it required some way of simulating the game, scoring the outcomes. And the idea with MuZero was that, well, we already have a deep neural network, right? These networks can learn a lot of things. So why not teach it to predict the future of the environment, the future of the world?

Lessons for today’s agents: search + learning at scale

Julian Schrittwieser [33:23] Why not make the model be able to learn for itself what is going to happen after each action it takes?

Matt Turck [33:58] After that, you also applied this to code and math. So that was AlphaCode and AlphaTensor. So zooming out a little bit, that evolution of reinforcement learning in games and then code and then math, what did you learn about the general power of search and learning that is today relevant in modern agentic AI systems? How did that whole body of work translate to what you are doing today?

Julian Schrittwieser [34:24] Games are a really good sandbox to learn very quickly about a lot of the reinforcement learning science: the algorithms that work well, the kind of problems that we encounter, even from a technical point of view. How do we build a learning system that spans many data centers, uses tens of thousands of machines? Because games are very clean sandboxes, very clean environments, so we can make many good experiments. And then now that we have a much more general model, right, the language models can do almost any task, but they're much more complicated and much slower to experiment with.

Do LLMs already have implicit world models?

Julian Schrittwieser [34:58] We can apply those same lessons of, we know how to build a really robust reinforcement learning infrastructure, and now we can build the same one for language models. We know if you do this kind of RL, then the model will learn how to exploit the reward. And so we can apply the same lessons, the same mitigation techniques to the language models.

Matt Turck [35:23] If I understand correctly, I think MuZero had a learned world model. So basically, rehearse the future, for lack of a better expression. Do modern LLM agents have anything like that? Do they have an internal world model that lets them preview actions before they commit?

Julian Schrittwieser [35:46] I think yes. I would say that language models have not an explicit world model, but they do have an implicit model of the world because, to be able to predict what is the next likely word in this sentence, how is this paragraph going to continue, they need to internally model what is the state of the world that makes this person say that thing. And so it's actually somewhat similar to MuZero in the sense that MuZero also only had an implicit world model.

Julian Schrittwieser [36:10] It was never trained to predict what does the screen actually look like if you take an action. It was also only trained to implicitly predict: if I take this action, what is the next action I should take? Or is it going to be good or bad for me? So in both those cases, you have an implicit representation of the world in your model that you can use to make predictions, but you're not actually reconstructing the full state of the world because reconstructing the full state of the world, that can be very expensive and complex.

Julian Schrittwieser [36:53] If you think about super-high-resolution video, audio signals, it's a very large amount of data that probably you don't actually need. If you think of human attention, we are only aware of a very small subset of what's actually going on all around us all the time because that's the most relevant information that we actually need to make decisions.

Matt Turck [37:24] And that goes back to the prior discussion about pre-training. So the reason why pre-training and RL work well together is that you have that world model that's implicitly embedded into the corpus. Although the argument against it is that it's what humans think the world model is, as embodied by language, versus what the world model actually is. And that's my understanding of the debate.

Julian Schrittwieser [37:28] I mean, for the debate, I think different people have different points of view, so I don't want to speak for anybody.

Matt Turck [37:29] Yes.

Julian Schrittwieser [37:57] But yes, I think pre-training on this rich knowledge gives you some representation of the world already so that when you actually start to act and interact with the world, you can very quickly make meaningful decisions, meaningful actions. I like to think of it in a similar way. If you look at many animals, when they are born, they very quickly know how to move, how to run even, right? If you look at gazelles, for example, in the savanna. In a way that is like, clearly they did not have time to really learn this from scratch, right?

Julian Schrittwieser [38:18] A few minutes or hours. And in their cases, they did not do pre-training, but they have some evolutionarily encoded structure in their brain. Because clearly it is very beneficial to have some sort of knowledge to make your learning more efficient.

Matt Turck [38:27] Yeah. Just RL in nature would lead to not so good results. Like if you're a gazelle and you have to A/B test whether to run towards the lion or away from the lion.

Julian Schrittwieser [38:50] Exactly. It's like thousands of generations of gazelles acquired this knowledge over time. It was encoded in their genes and their brain structure in some way. And then you get to start on top of that. I think the main challenge or the main thing you need to watch out for is that you don't over-encode or you don't restrict your search space too much. If your pre-training, if your prior knowledge prevents you from exploring something that might be the correct course of action, that will be bad.

Why RL on LLMs took time (stability, feedback loops)

Julian Schrittwieser [39:02] So there is some danger there you have to be aware of.

Matt Turck [39:34] So this general idea of making pre-training and RL work together in modern AI systems seems to be the big idea or topic of 2025, although of course I know it's been years in the making. Why did it take so long? It feels like RL progressed in its own direction and then pre-training worked in its own direction, and those were slightly separate. Why did it take so long to put them together? Is that just purely practical and economic, or anything else?

Julian Schrittwieser [40:01] Scaling up the language models to the massive degree that we scaled them up took a lot of effort on its own. And from a science point of view, from an engineering point of view, pre-training and supervised training are more stable and easier to debug because you don't have this feedback cycle. You basically have a fixed target and you're trying to learn this target. And so then you can focus on, is my training working? And is my infrastructure working? And then, does it scale?

Julian Schrittwieser [40:15] Does it fall over? Versus if you compare to RL, in RL you have this feedback cycle of, I learned something and then I use that to generate my new training data. And then I learn from that training data.

Matt Turck [40:15] Right.

Julian Schrittwieser [40:43] And now if something is not working, it's very hard to figure out where in this cycle your problem is coming from. Maybe your training update was bad, and that's why you suddenly started behaving badly. Or maybe the way you decided, the way you select actions to behave, is not correct. And so you generate bad training data, and that's what messed up everything. So it's just much more complicated to get working correctly. And so I think it makes a lot of sense to first scale up the pre-training, the architectures, figure out something that works pretty well, especially if you can already get pretty far by some fine-tuning, some prompting.

Julian Schrittwieser [41:17] And then when it's clear that these models are really general, they are really useful, and we have them in a pretty stable state, then you can ramp up RL and take them even further. Even in our own work, if you look at AlphaGo, AlphaZero, we always followed a similar split as well, where we first set up the architecture of the network, the training using fixed supervised data. And only when we had that working really reliably, only then did we do the full RL loop and the full training.

Compute & scaling for RL — what we see so far

Julian Schrittwieser [41:43] Just because debugging all of it at the same time, you're just setting yourself up for failure. It's really useful to be able to isolate the component and say, I have known-good data over here. I have a known-good target there. If the thing in between is not working, I can isolate it. And then we can isolate all parts of the system.

Matt Turck [41:57] How compute-intensive is it to scale RL? And are there scaling laws for RL the same way you do in pre-training?

Julian Schrittwieser [42:24] There's less published literature about it. But I think if you look at all the RL literature over time, we see very similar returns on compute in pre-training and in RL, where we can invest exponentially more compute in RL and keep getting benefits. There's going to be some interesting research to come to figure out what are the trade-offs between pre-training and RL compute. We don't know what should be the split for a big model, for example. Could it be 50-50? Should it be, like, 1 to 10? Which way should it be 1 to 10?

Rewards frontier: human prefs, rubrics, RLVR, process rewards

Julian Schrittwieser [42:36] So I think that's going to be extremely interesting. But so far, yeah, we definitely see good returns on both.

Matt Turck [43:09] What's the latest state of the art or thinking in the field of rewards? So in what you described for AlphaZero, AlphaGo, that was basically win-loss as a reward. Then it sort of feels like we went into kind of fuzzy human matching: this is good, this is not good. And now that we expand, as per the above, into more general fields where it's sort of unclear whether you win or lose, how does that work? What parts of the evolution are you working on?

Matt Turck [43:15] Are you excited about?

Julian Schrittwieser [43:52] Personally, I don't work that much on reward modeling. I mostly work on sort of reasoning, planning, search, compute, ways of making the model smarter by spending more computation. Yeah, thinking about rewards, I think the reinforcement learning process per se doesn't really care where the reward comes from. The algorithms are very happy to use any source of reward, whether that's a human feedback signal, some automated signal from winning or losing the game, or passing a test, whether it's something more model-generated. For example, at Anthropic, we had this paper about Constitutional AI to have the model itself score whether you're following some guidelines.

Julian Schrittwieser [44:07] So it can be very flexible as to what kind of reward we follow.

Matt Turck [44:18] But RLVR, all those things are at this stage stuff that you see commonly used. Any thoughts?

RL training data & the “flywheel” (and why quality matters)

Julian Schrittwieser [44:37] Yeah, I think we're seeing a huge mix of rewards and environments. And I think people are working very hard on figuring out what are the best reward sources, and how do we scale it up, and how do we get more rewards, more reliable rewards. That will be one of the key ingredients in scaling up RL further.

Matt Turck [45:03] And so, switching from rewards, what is the latest thinking in terms of training data for RL? Again, following the evolution from AlphaGo, where it used to be human data and then self-play. How does that work? Where does the data come from, and what kind of data works best to train modern RL?

Julian Schrittwieser [45:28] Yeah, I guess a great thing about RL is that the data is generated by your model itself. So the smarter our models become, the better RL data we can generate, the more interesting and complex tasks they can solve, which then gives us more and more data that we can train on. Because the more complex the task, the longer it takes to solve the task, the more data it generates that we can then use for training. I think part of the challenge is to find tasks that are really representative of what people actually want to do with the model.

Julian Schrittwieser [45:52] Because now language models are so general, people are using them for so many different things. There's more and more of a challenge that we need to cover as many of those as possible in our RL to make sure that the model is actually able to do this diverse set of tasks.

Matt Turck [45:59] What matters more for training data? Is it quality? Is it quantity? Is it recency?

Julian Schrittwieser [46:30] I think that's a very interesting question that maybe doesn't have a super clear answer yet, or maybe there's still interesting research to be done. I think we've seen papers arguing for different things, or we've seen different benefits. Clearly, we see in pre-training, as we scale up the data, we can keep improving, but we've also seen very interesting fine-tuning results, papers published where, with a very small amount of examples, you can teach the model how to do an interesting skill. And I think we don't have any good scaling laws yet that tell us the trade-off, especially, I think, because it's very hard to measure what is the quality of a data point, right?

Julian Schrittwieser [47:07] How good is this example compared to this other example? Without being able to measure this, it's very hard to quantify the trade-off in any way. I think intuitively, it's definitely true that if you have bad data, RL doesn't work that well. And if you have very high-quality data, it becomes much more stable. For example, I think that was very clear in the AlphaZero days, where AlphaZero spends a lot of computation. It does a lot of planning and search to decide which move to take.

Julian Schrittwieser [47:25] And so that generates very high-quality data to train on, which then resulted in RL training that was incredibly stable. So you can run it across continents, take a long time to generate the data, and then train on it. And it's very robust, versus in modern RL with language models, the difference in how good the model is and what data it generates that we then train on is not so large because we more directly sample from the model and then train on it, which then results in reinforcement learning that is less stable.

RL & Agents 101 — why RL unlocks robustness

Julian Schrittwieser [48:02] And so one direction of scaling RL and making it more stable is by improving this, by, for example, putting more reasoning into your language model to generate much more high-quality training data that can then give us training that is much more stable and that we can scale up much more easily.

Matt Turck [48:35] I'd love to spend a little bit of time now on the general topic of RL and agents. So, the famous agentic AI that everybody's been talking about breathlessly for the last year. So, for people listening, and as often in an effort to make this broadly accessible to a group of people in tech, could you drive home the intersection and overlap between RL and agents? Does RL power agents? How does that work?

Julian Schrittwieser [48:43] Yeah, so I guess maybe first let's take a step back on what do we actually mean by agent as compared to a general language model, right?

Matt Turck [48:47] The second most debated question after AGI is: what is an agent?

Julian Schrittwieser [49:09] Yes. I guess, yeah, for our purposes, let's just say that an agent is an AI that can act on its own. Maybe take some actions on a computer, save some files, edit some files, send an email, whatever you want. But the main characteristic is that it doesn't have to interact with the user all the time. It can do things on its own. The reason why RL is very important for this actually connects back to pre-training, because our pre-training data is not very agent-like.

Julian Schrittwieser [49:46] If you think of the pre-training data, there are websites and books and all kinds of written text that has a lot of information, but it doesn't have a lot of actions. It doesn't really capture how humans actually interact with the world. So if you take a raw pre-trained model, it's not a very good agent. Maybe you can prompt it a bit and sort of push it in the right direction, but it's not going to be very good at interacting. And especially it's not going to be very good at correcting for its own errors, because the pre-training data has no examples at all of how our agent is going to fail.

Julian Schrittwieser [50:27] And that's exactly where reinforcement learning comes in, because in RL, we can take our agent, let it interact with the environment, and then directly train on that interaction. So, for example, if the agent did well, we can reinforce those actions. And if the agent did badly, we can push it away from those actions. And if the agent sort of did badly at the beginning, but then recovered and managed to do well, then we can also reinforce that recovery. And so that's super important because it allows the agent to actually learn from its own distribution of behavior.

Should builders use RL-as-a-service? Or just tools + prompts?

Julian Schrittwieser [50:52] And that just makes it much more robust because now it doesn't have to generalize to something it has never seen before. It can actually learn on the actual problem that it's trying to solve. And that's why RL is really unlocking so many agentic capabilities now.

Matt Turck [51:10] If I'm an AI builder today building an AI app, and I build it on top of Anthropic's, or whatever model, it's going to come with some of this sort of batteries included. But as a builder on top, do I need to do my own RL? There is this emerging space of RL as a service where, for this task or that task, I build on top of a general model that sort of offers the ability to do RL, or can I do a lot of that just through prompts or maybe supervised fine-tuning first?

Julian Schrittwieser [52:00] I think nowadays, with the capabilities of top Anthropic Claude models, top OpenAI GPT models, you don't need to do any fine-tuning. You can take the model as is, write your own tools, your own harness, and benefit from that agentic training, because doing good agentic fine-tuning is actually very hard. And so it's quite hard to do better than the top frontier models that you might get. But on the contrary, coming up with good tools and a good representation of your task makes a huge difference.

What’s missing for dependable agents (capability vs engineering)

Julian Schrittwieser [52:18] So depending on how you express your problem for the model can make it way harder or way easier. And so you can get a lot of mileage out of that.

Matt Turck [52:39] What's currently missing to achieve the big dream of agentic AI? Is it model capabilities at the core, or is it sort of like boring engineering around reliability, tool use, safety? What needs to happen?

Julian Schrittwieser [53:12] I think there's basically improvements needed around the whole space. Make the model better able to correct its own errors, make the model better able to continue going for long times without getting distracted, make the model just smarter in general, maybe make the model faster. There's basically a whole set of things that we know that we can improve. There's probably not one individual blocker, and that's why we will continue to see smooth incremental progress over model releases. But sort of given how many things we know there are that we can do better on and improve...

Julian Schrittwieser [53:41] Yeah, I'm quite excited about where models are going to end up. I think that's actually one of the reasons why AI is a very fun field, is that there are so many low-hanging fruits that you can do much better on, but already the current models are so good that it's very fun to work on it. It's like, oh, I can fix this thing. It'll be even better. Versus if you're in a place where everything has already been solved and it's really hard to figure out how to make it better, it's a very different story.

Evals & Goodhart — internal vs external benchmarks

Matt Turck [54:14] Let's spend a minute on evals. We touched upon this a little bit, but just to give it some proper space. So there was, in your blog post that we talked about at the very beginning of this conversation, this concept of external benchmarks, and then you quoted this piece, Goodhart's Law. First of all, what is Goodhart's Law? And then how should labs compare results so that it doesn't end up with this leaderboard theater that we've seen a little bit in the last couple of years?

Julian Schrittwieser [54:53] Goodhart's Law basically says that any measure that becomes a target stops being a good measure. You can think of that intuitively: if you start paying, for example, programmers based on how many lines of code they write, suddenly they will discover many ways to add more lines of comments, which is completely useless. And this is a very general effect. If you give people an incentive that they should optimize, they will try very hard.

Matt Turck [54:54] Yes.

Julian Schrittwieser [55:16] And we also see this with language model benchmarks. Of course, people want to get promoted. They want to launch their model. So any benchmark that is too easily measured or that has a lot of attention on it, people will optimize very hard for it, which means that probably the model will look very good at that benchmark. But if you then use it for your own task, you might get different performance. You asked, what do we do about this?

Julian Schrittwieser [55:23] It's very hard to prevent people from optimizing on the benchmark.

Matt Turck [55:24] Right.

Julian Schrittwieser [55:53] So one possibility is just to periodically create completely new held-out benchmarks that nobody has seen before. And that gives you a fairly good estimate of model performance. I know, for example, a lot of researchers have their own toy problems that they use to test all the models precisely for that reason. So that this is a set of problems that nobody has seen, you have a pretty good guess that it's going to give you an unbiased estimate. If you're an individual or if you're a company trying to decide which model to use, it's probably something similar.

Julian Schrittwieser [56:13] Just make your own internal benchmark that really represents what you care about and then measure on that. And I think that's likely to be the most objective, most accurate way of measuring.

Matt Turck [56:31] Internally, what does that look like at a place like Anthropic or previously DeepMind? I mean, I know there are teams that are focused on evals. How do you think about what works, what doesn't, in terms of internal evals?

Julian Schrittwieser [56:53] It definitely used to be easier to have good evals. Five years ago, the tasks we were doing, I think it was easier to measure model performance. I think nowadays it's much more difficult, and I think we try not to over-rely on evals so much because it's quite hard, for example, to measure how good this model really is at writing code. I think it's one of the big unsolved, or very important, problems in the field: making really good evals that are both cheap to run, reliable, and accurate. Because it's easy-ish to make an eval that ticks one of those, but to get all three is quite hard.

Mechanistic interpretability & “Golden Gate Claude”

Julian Schrittwieser [57:35] For example, at the beginning, we were talking about OpenAI's GDPval. And that one is very accurate and unbiased, but it's very expensive to run because what it actually involves is taking human experts, having them do the task, and then comparing the model task to the experts and rating it with multiple people. So it's very accurate, but it's extremely expensive to do.

Matt Turck [58:15] Related to that topic of evals, what's the latest in terms of our ability, or I should say your ability, to truly understand how models work? So, the general field of mechanistic interpretability. You alluded to the fact earlier that RL, if I understood correctly, sometimes makes it a bit harder because it does things occasionally in a more inscrutable way. My words, maybe not yours. So what is the latest? And indeed, does RL make things harder or easier?

Julian Schrittwieser [58:39] So what I meant before is that debugging RL in general, completely unrelated to interpretability, is harder because there are more moving parts. But it is also true that if you're not careful with RL, you can make interpretability harder. For example, one common thing with modern models is they do reasoning with the chain of thought. You could look at the chain of thought to see what are the model's internal thoughts. And then you could also have a thought that, maybe I should use that as a reward signal in RL.

Matt Turck [58:45] Mm-hmm.

Julian Schrittwieser [59:08] Punish the model if it thinks the wrong thing. But then suddenly you completely destroyed your interpretability angle. So you sort of have to be careful that you don't do RL on the signals that you actually want to use to interpret what the model is thinking or doing. That said, I think there are some extremely exciting interpretability things happening, including mechanistic interpretability. I think actually last year, I think before Anthropic, maybe even, there was a super cool Golden Gate Claude model where they found the neurons in Claude that were responsible for the Golden Gate concept and then modified them to make a version of Claude that really loved the Golden Gate Bridge in San Francisco.

Julian Schrittwieser [59:49] And so that's a really vivid example of, oh, we really understand what's happening in this model. And what better way is there to verify that understanding than actually changing the behavior of the model? And so I think that's a super important direction for safety. As the models get smarter, we really need to be able to understand: what is the model thinking internally? What values does it have? Is it lying to us? Is it actually genuinely following the instructions?

Safety & alignment at Anthropic — how it shows up in practice

Julian Schrittwieser [1:00:03] And so I think definitely extremely important area to invest in and work in. I think especially if people are interested in working in AI or doing AI research, I think interpretability is a great area to get into.

Matt Turck [1:00:35] Yeah, perfect segue for the last part of this conversation. I'd love to zoom out and talk about the impact of AI. So if we think that we are on the exponential and that things are going to only accelerate from here, what does that mean? And certainly safety and alignment, which is a core value at Anthropic, hopefully in other parts of the field as well. But Anthropic is particularly vocal about safety and alignment, let's say. How does that actually manifest?

Matt Turck [1:00:56] So we just talked about interpretability. For people who are concerned that this is going too fast and that we collectively are creating a monster, can you give us a glimpse into the kind of work that is done for alignment and safety at a place like Anthropic?

Julian Schrittwieser [1:01:25] Yeah, I think the focus on safety and alignment pervades all of Anthropic, and there's very rigorous processes. When we train a model, whenever we want to release a model, both to analyze the capabilities of the model, verify the alignment of the model, ensure that it does not do harmful things on its own, ensure that it does not enable malicious users to do harmful things. And to the point where if we are unsure about the safety of a model, we will delay the launch until we're sufficiently sure that it is actually harmless.

Julian Schrittwieser [1:02:07] We will not launch and release a model, which may, I guess, show that people clearly take safety much more seriously than any financial return or revenue. I think also in terms of research and resources, the teams working on safety and interpretability are a big focus of the company, which gives me a lot of confidence that we actually care about this and put a lot of effort into it.

Matt Turck [1:02:39] And at a more technical level, and to tie back to an earlier part of the conversation around when we were discussing the pre-training and safety, so is safety and alignment an RL problem? And by that I mean the beauty of having pre-training is that you import that world model, as we were discussing. But arguably you also import into your brain a lot of bad stuff if you collect data from the internet. As we know, there's good things, but also a lot of toxic content.

Matt Turck [1:02:52] So is alignment largely using RL to get rid of the bad stuff that is built into the pre-training?

Julian Schrittwieser [1:03:21] We can definitely use RL to shape the model behavior and ensure that, for example, given bad input, it sort of behaves safely or knows that it can refuse or is robust to attempts to prompt-hack the model. Yeah, I wouldn't view alignment as just, like, an RL problem. I think it sort of goes throughout the whole stack. You might, for example, filter the pre-training data in some way. You might, after training, have classifiers that look at the model, monitor the model behavior to ensure that it is actually aligned.

Jobs: human–AI complementarity (comparative advantage)

Julian Schrittwieser [1:03:49] When you write a system prompt for the model that you use, you might put safety guidelines in there. So I think safety alignment really pervades the whole of research and the whole of product and deployment. It's not just isolated into any one part.

Matt Turck [1:04:20] And then another super interesting topic in the same vein of the impact of AI is obviously the discussion around jobs. So if, as per the GDPval discussion, the agents are becoming just as good or better than humans, obviously, what does that mean for all of us in terms of our jobs? What have you learned after the experience of AlphaZero, AlphaGo, that could give us a glimpse into what may happen once we all have super powerful agents do our jobs?

Julian Schrittwieser [1:04:48] So I think the first thing that we didn't talk about yet so far is that artificial intelligence is quite—I mean, this may sound a bit simplistic, but it's quite different than human intelligence. So we can see that, right, that the model may be much better than us on some tasks, like calculation, obviously, and much worse than us at other tasks. So I don't think it is at all going to be any one-for-one replacement.

Matt Turck [1:04:48] Right.

Julian Schrittwieser [1:05:03] It's going to be much more complementary. The model is really good at something that maybe I really don't like doing, or I'm not interested in, or I'm very bad at. And then I'm much better than the model at some other part. And so I think it's going to be like a gradual process of we're all going to incrementally start using models more and more to improve our own productivity rather than have a model that one-for-one is able to do exactly the set of things we can do.

Julian Schrittwieser [1:05:44] For example, I use Claude all the time to refactor code or maybe write some frontend code that I don't want to write. At the same time, there's other parts where I'm clearly much better at coding than Claude still. So there is a synergy of using the best, most productive skills. I think economists call it comparative advantage. But there is this long process of building both to sort of improve our productivity incrementally. And I think that process is going to give us some time to figure out politically and economically how do we want to benefit from this massive productivity increase.

Julian Schrittwieser [1:06:21] Even independently from AI, the promise of technology has long been that, oh, we're going to be all so productive, so wealthy that we need to work much less. Yet mysteriously, we all have like 40-hour workweeks for decades. And so I think it's much more like a political, social problem of figuring out how do we actually benefit from all these improvements and bring the increases in wealth and productivity to everybody. And it's much less a technological problem, which also means that we can't really solve it with technology.

Inequality, policy, and the case for 10× productivity → abundance

Julian Schrittwieser [1:06:33] We have to solve it at a democratic, political level. How do we spread these benefits?

Matt Turck [1:06:55] Do you think that increases inequality? So, as you think about the impact of AlphaGo and MuZero, what happened to the top Go players and what happened to the top chess players? Did they disappear, or did they get enhanced and better?

Julian Schrittwieser [1:07:22] Yeah, I think at least in the case of chess and Go, there has been more interest, and it has become much easier for people to study how to play Go, how to play chess, because now you don't need to find an expert tutor. They can practice on their own, spend a lot of time. I guess chess streamers are very popular on Twitch right now. And similarly, a lot of students are using language models to study. I think also for coding, Claude Code, these agents, they raise the bar of what anybody who has an idea can accomplish on their own.

Julian Schrittwieser [1:08:06] I think the larger picture, whether it increases or decreases inequality, is quite hard to forecast. It both sort of raises the floor of what any person can accomplish, but it also gives very productive people an ability to be even more productive. It's possible that we see quite a difference between countries depending on the taxation and social redistributive system that they have, in whether inequality increases or decreases, for example. Overall, I'm quite excited that it is very much non-zero-sum. It very much increases the total wealth available in society.

Julian Schrittwieser [1:08:40] I think if you think about progress, if you think about prosperity, that is the most important thing. Redistributing the pie is kind of a loser's game. To get more wealthy, we really need to grow the pie. If you think of the agricultural revolution, the industrial revolution, the reason why we have much better lives nowadays is because we are so much more productive, we have so much more wealth. And so that's the key step we want to unlock. If we manage to make everybody in society 10 times more productive, what kind of abundance can we achieve?

Julian Schrittwieser [1:09:15] I think that's the key question, right? What advances does that unlock in medicine? Curing diseases, halting aging. What does it unlock in terms of energy? We have a climate crisis, we need more energy to sustain our lifestyle. What advances in material science can we have? All of those are basically bottlenecked on how much intelligence we have access to and how we can apply it. So, I'm incredibly optimistic about what we'll be able to unlock in the next five years.

Closing thoughts

Julian Schrittwieser [1:09:24] I think we can go extremely far.

Matt Turck [1:09:32] Well, that feels like a wonderful place to leave it. Thank you so much, Julian. This was absolutely fantastic. Thank you for spending time with us.

Julian Schrittwieser [1:09:35] Yeah, thank you for all the exciting questions and giving me the time.

Matt Turck [1:09:56] Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you on the next episode.