Open Source AI Strikes Back — Inside Ai2’s OLMo 3 ‘Thinking"
The MAD Podcast with Matt Turck · with Guest
Guest is the Guest at Ai2. We cover why OLMo 3 releases its data and intermediate checkpoints alongside model weights, why Chinese labs use open models to gain U.S. mindshare, and why small reasoning models get most of their gains from distilled training traces while RL infrastructure remains technically punishing.
Chapters
- 2:07 — What “base models” really are (and why they matter)
- 5:51 — Dolma 3: the data behind Olmo 3
- 8:06 — Performance vs Qwen, Gemma, DeepSeek
- 10:28 — What true open source means (and why it’s rare)
- 12:51 — Intermediate checkpoints, transparency, and why AI2 publishes everything
- 16:37 — Why Qwen is everywhere (including U.S. startups)
- 18:31 — Why Chinese labs go open source (and why U.S. labs don’t)
- 20:28 — Inside ATOM: the U.S. response to China’s model surge
- 22:13 — The rise of “thinking models” and inference-time scaling
- 35:58 — The full Olmo pipeline, explained simply
- 46:52 — Pre-training: data, scale, and avoiding catastrophic spikes
- 50:27 — Mid-training (tail patching) and avoiding test leakage
- 52:06 — Why long-context training matters
- 55:28 — SFT: building the foundation for reasoning
- 1:04:53 — Preference tuning & why DPO still works
- 1:10:51 — The hard part: RLVR, long reasoning chains, and infrastructure pain
- 1:13:59 — Why RL is so technically brutal
- 1:18:17 — Complexity tax vs AGI hype
- 1:21:58 — How everyone can contribute to the future of AI
- 1:27:26 — Closing thoughts
Transcript
What “base models” really are (and why they matter)
Matt Turck [1:46] Guys, welcome to the pod. Big announcement today and a big day for open-source AI. Walk us through what it is that you're releasing today. Thanks for having us. Yeah, we're launching the OLMo 3 family today. So this is our latest family of open-source models. We have a 7B model, a 32B model. We have models that can think, models that can follow instructions and use tools.
Matt Turck [2:21] And just like every single model that we released before, we're not just releasing the final models. We are releasing the entire recipe we followed to get this model: the data, the intermediate states, the evaluation frameworks, all the details, all the bits that people need to know to make models like OLMo. Specifically, there's OLMo 3 Base 7B and 32B. So what are those? Probably, we have, say, five flagship checkpoints that we're putting out. Two of them are base models.
Matt Turck [2:48] That means these are models before they get trained to respond to user instructions. So these are really good for folks who want to take the sort of bulk of our compute that we spend in pre-training these models, and then they want to customize them for their use cases. So these are a 7B model—there's a smaller one that's more efficient that takes about one GPU to fine-tune for a use case—and then there's a larger 32B that takes about one box of eight GPUs to fine-tune.
Matt Turck [3:21] And then on top of that, we have our fine-tuned, our post-trained models for various use cases. So there are a couple of models that are thinking models. So there's OLMo 7B Think and OLMo 32B Think. These are models that, just like a lot of the reasoner or pro models out there, can spend compute power at inference time to think through a problem and solve it, and then give you an answer at the end. And also, we are releasing an OLMo 7B Instruct model.
Matt Turck [3:36] This is a more immediate model that gives you faster responses. So it's really good for bulk data processing or use cases where you want to have low latency in your responses.
Guest [3:56] I want to add more color to these things. I think Luca's underselling the base model. We're going to talk more about this, but over this year, a lot more people have been releasing open models, especially large open models. But some people are starting to not release base models. We have a bunch of DeepSeek-sized giant MoE base models and a bunch of small base models. But, for example, Qwen 3, which everyone accepts as a research standard and an industry standard, they don't have this 32B base model.
Guest [4:23] OLMo 3 32B is still the best base model. The upside is that we have all the data, so people can actually do some sort of continued pre-training and hopefully make it a bit easier to modify and understand the behavior. That's exciting for us, though. The actual potentially best-in-class thing. It's also a fully open thing, which is not something we get to say a lot at Ai2. Sometimes it's like, oh, we replicated this and now you can do it yourself.
Guest [4:47] This is actually a good thing. And then 7B models, which Luca was saying, used to be this huge industry standard where there's just so many of them. It's still a standard size category, but there aren't quite as many models there as there used to be. Especially the instruct models are less common, and this is up there with one of the best in the world there at that size category again. OLMo 1B is one of the most used models on Hugging Face of all time, and this should be better than OLMo 1B, I hope.
Guest [5:03] I hope that holds up for people. We can release more and fix it. But that's just trying to give—we might not be at this frontier scale, but these are things that are still widely used in the world. Then it's the first fully open reasoning model where we show doing RL on base models and distilling from bigger thinking models and all these things that people have seen a ton of times throughout the year. I think another thinking model is like, ooh, what is this one for?
Guest [5:38] But we have all the data and we show people what to use it for. So I think a lot of times with our, especially open post-training, it's just like the datasets become a standard. So it's like our Tulu 3 dataset from last year, which we use for OLMo 2, is in, like, the Thinking Machines Lab Tinker API. And we want people to use this data, modify it how they need to, and look at the different training stages.
Dolma 3: the data behind Olmo 3
Matt Turck [6:18] So, to the data point, talk about Dolma 3. Dolma 3 is the data that we use in pre-training for OLMo 3. So it's what we use to create the base model. It's really three parts. There's the sort of pre-training pool. This is like a pool of about 10 trillion tokens from which we have, like, an algorithm, also fully open source, to sample about 6 trillion tokens that we use during training. And we have kind of new techniques there.
Matt Turck [6:34] It's kind of interesting. We do this technique where instead of repeating at random documents to get more training tokens, we intelligently repeat the tokens that have the most value.
Guest [6:35] So we have that part.
Matt Turck [6:57] There's a smaller subset that we use during this mid-training phase. So this is a more focused dataset with a lot of math, high-quality code, sort of knowledge tidbits that you want a model to pick up. And then finally, we have a set of documents that are particularly useful to make models able to work with long context. Really excited about this one because historically, of the data that is available openly out there for people to build their language models, you don't have a lot of long-document data.
Matt Turck [7:41] So these are documents that we crawl ourselves. They're PDFs, mostly science PDFs. They're openly available on the internet to crawl. We have a pipeline, but it's also open source. Everything's open source to turn them into plain text. And of those, we have, instead of web pages that are kind of short—like 95% of web pages are below 3,000 tokens—these are quite long. We have about 600 billion tokens that are longer than 8,000 tokens. So these are really good for people to develop other ways for models to understand very long inputs, which is typically something that people are not able to do today in the open unless they are a big lab and they can acquire data that is long enough to.
Performance vs Qwen, Gemma, DeepSeek
Matt Turck [8:29] Thank you. You alluded to some of this, but talk about performance and efficiency. Performance, it's very hard to measure performance of a base model. So for the instruct and the thinking, Nathan will have more info about comparative benchmarks, but the base model is really good. Qwen 3 or Gemma 3, certain capabilities they have, maybe a little bit better on some capabilities, better on some others, but sort of where it's there in the ballpark. Absolute performance of a base model doesn't matter so much as an instruct model at the end.
Matt Turck [8:46] You want to be in the right band where the model is capable enough that then your post-training team can do magic on the checkpoint and make it really, really good.
Guest [9:20] I would say in post-training, we're the best models that don't start with Qwen 3. We're reasonable to say that they are comparable to Qwen 3. On some benchmarks, we beat them, but on some benchmarks, they're way ahead. I think a lot of people are like, we don't know what Qwen 3 puts exactly in the training data, so we don't know if some benchmarks they benchmarked MATH a little harder than we did. I mean, we try to hill climb on benchmarks to make our model good.
Guest [9:41] I think there's always some level of this. But in that, it's in the same ballpark, plenty of things. We're hoping that there's use cases where people that use Qwen 3 8B or 32B are willing to switch over and get some value out of this and maybe modify it to their own use cases. But it's like, Qwen also releases great models. So it's like a never-ending uphill battle that motivates you to do better, to try to get close and compete with what they're doing.
Guest [10:07] It's like they released these Qwen 3 VL, their vision models, and on text-only benchmarks, it's way better than the models they released in April. So it's like, okay, that's the new baseline, and most people don't know about it because they think it's just a vision model, but it's actually a much better text model. The bar's always rising. But at the 7B scale, NVIDIA had Nemotron Nano V2, which is a 9B hybrid model, which I think is almost equivalent to our 7B model.
Guest [10:27] These are good models. There's not that many of them that are in these size bands that are really strong. Happy to be there and happy to point out other people are doing great work here. It's not like we can ignore Qwen. That's a losing strategy.
What true open source means (and why it’s rare)
Matt Turck [10:57] Luca, just to drive it home, the concept of open source in AI, there's different flavors of it. Walk us through what that means and where you guys are at. That's always a topic that gets sort of overlooked a little bit in discussion. But yeah, when it comes to models, there are different levels of what people consider open source. The majority of models that get released, I think the best term to describe them is open weights. Your Qwen, your Gemma, your Llama, Kimi—what gets released is a set of weights that correspond either to the final state of a model, that's the most common, or maybe the final state of the instruct model, final state of the base model.
Matt Turck [11:41] And there are plenty of cases where that's enough, and you can build great software on top of it. There is an equivalent large set of cases, from research to application, where that's just not enough. You want to have intermediate states of the model so that you can customize it better. You want to have access to the data so you can maybe redo a step of the training while infusing your own data. You might want to have access to the pre-training data because you have this incredible research project that's going to change how we think about language models, but you need to know what a language model was trained on.
Matt Turck [12:27] So we want to support those use cases. So when it comes to OLMo, if we can release it, we will release it. We can't release our GPUs out to the world. That's not how it works. But when it comes to the data, the intermediate checkpoints, the benchmarks, the software—anything we can, we'll put it out. If people ask, "Hey, you described this part of your pipeline, but you haven't put it out," we'll release that part as well.
Intermediate checkpoints, transparency, and why AI2 publishes everything
Guest [12:51] We've always got questions about intermediate checkpoints during SFT or other fine-tuning stages. Now we have intermediate checkpoints during our supervised fine-tuning for reasoning and for instruct, and then also for our multi-day RL rounds. At the end of these, we have intermediate checkpoints. So people that are looking—a lot of people like to understand checkpoints and do research on them, but don't have compute to train. And now it's like, okay, this is all there.
Matt Turck [13:19] Before we dive into the specifics, I'd love to take a step back. It's been a very intense year in the world of open-source AI. The DeepSeek moment feels like it was three years ago, but in reality, that was at the end of January, so 10 months ago, and a lot has happened since. Nathan, could you help us maybe recap the key events of 2025 for people to understand what's happened?
Guest [13:43] Yeah, if I try to make a list of actual models, I'm going to forget some because there are so many that are notable. I think starting with DeepSeek, as you mentioned, is definitely the important thing. And then if you talk to people building models in China, a lot of the consensus is that DeepSeek showed us that AI could be a big deal and that a lot of these companies were like, oh, we should do what they did. So there's just a ton of labs that have popped up over the year.
Guest [14:16] Z.ai and Kimi Moonshot had already existed. These really stepped up to be much more known names, especially if you're following Western, SF-centric discourse. These are things that people are using and talking about, which is a kind of big change. But there's just this huge mass of models coming from China. You have everything like Ant Group is releasing trillion-parameter MoEs with really strong benchmarks. Meituan, which is the Chinese equivalent of DoorDash, which is just another big tech company in China. The standard way of developing language models has become to release them openly, and that whole ecosystem is going forward with this, figuring this out.
Guest [14:48] When, at the same time, there was a big change in leadership at Meta, and Llama's future is less known, which was really the paradigmatic definition of open-source AI. That line of thought just ended. So there's this big vacuum of influence which has been absorbed by the likes of Qwen, DeepSeek, Kimi Moonshot, in terms of who's trying to build things with open models. That's a big shift. I think there's a lot of discussion within the U.S. that there's good reason that we should own, we should at least have influence over the whole technological stack, and that includes open models.
Guest [15:22] Realistically, it's the big tech companies in the U.S. that'll capture the downstream value of that from having the researchers be in close proximity and speaking the same language and used to the infrastructure. I think this is something that we've seen for decades in the tech industry, so I don't think I need to explain it that much. There are people that are really starting to wake up to this. I think June, July is when the Chinese model providers were really becoming like, you could not ignore them.
Guest [15:51] That's when we had the Kimi K2 Instruct, GLM, and that's kind of just continuing now. So I think when we're recording and releasing this podcast, there's a lot of interest in, like, what are the U.S. companies going to do to respond to this? I know that NVIDIA is making a lot of noise here. They invested a lot of money in Reflection AI, and there are other players that are trying to get going. But urgency—and we don't have a lot of compute at Ai2—but if we can make a dent in this in some model sizes that people actually use, I think that—we focus on researchers.
Guest [16:18] I think dense models are great for researchers, so they take a little bit less compute and engineering resources to use. And that's—I do think that there's more U.S. participation. I mean, OpenAI has released some models, where they're just established in a different way.
Why Qwen is everywhere (including U.S. startups)
Matt Turck [16:55] And Qwen is widely, widely used in a way that people may not have completely realized, right? There was, as an anecdote, the example of Airbnb talking about using Qwen over ChatGPT a few weeks ago, but do you have any stats or anecdotal evidence on the usage of Qwen?
Guest [17:20] The other famous quote was a Martin Casado quote in The Economist where he said 80% of companies are building on Qwen. That has been corrected, where it's 80% of companies building with open models are using Qwen, which is like 16 to 24% of his portfolio, which is still a lot. It's a meaningful amount of people who are trying open models for things, and most of them are using Qwen. Then there's the likes of Cursor, who released their own model, Composer 2. It's suspected that it is built on a large Chinese MoE of some sort that was released openly.
Guest [17:50] There's some obvious tells of it switching to Chinese and things like this, but that is the sort of company that doesn't want to pre-train their own models but has immense value in specifying models for their use case that is just going to build on these great models. I think they would want more options to choose from as they try to sell into more markets. I think realistically, it's a thing where a lot of U.S. companies don't want to deploy Chinese models. I think currently a lot of the stated reasons are just unknown unknowns and things you can't prove.
Guest [18:10] You can't prove that the models aren't doing certain backdoors, where I'm fairly certain they definitely aren't now. But just because you can't prove it makes this kind of weird market dance, which is like, yes, these are stochastic things that are kind of amorphous. And it's like, I don't love being in the middle of this as a researcher, but it's like, I would like to just provide information and good things that people actually really want to use and leave all of the geopolitical and other messaging to people that have probably, realistically, way more on the line than I do.
Guest [18:30] Like, I don't know, we work in a nonprofit. I have my dog.
Why Chinese labs go open source (and why U.S. labs don’t)
Matt Turck [18:43] Why do you think this happened, that the ecosystems developed in this way, that the U.S. was very commercial, closed-source, and China very open-source?
Guest [18:57] Historically, the U.S. has a lot more willingness to pay for services. I hear anecdotes from people that know China a lot more than I do that are like, yeah, medium- to large-, billion-dollar-plus valuation companies in China will just pirate SaaS software. I know that sounds worse than it is, but it's just, I think the thing is that U.S. companies are used to paying for services, and that API model and paying for tokens has been proven as a very good—selling tokens is a good business in the U.S. right now.
Guest [19:20] I think there's a lot of debate over profitability, but the demand and usefulness of these tokens is high. So I have a lot of belief that there can be profitable businesses from selling tokens. Where I think that AI will be embedded in very different ways when it comes to Chinese companies, and I've talked to a few of these labs, and they're like, in order to sell into the U.S. market, they will not pay for—they've said this—U.S. companies will not pay for services.
Guest [19:58] So they don't expect enterprises to sign up for the Kimi coding plan en masse, but they're like, we have a chance that they'll use our models. And it's like, that is a practical way to influence and getting a piece of the sharing pie. And it's like, the people building these models in China know the same things about the different ecosystems. That's why I've enjoyed starting to talk to them. It's like, oh, these people, that's the same thing. They see the same constraints.
Guest [20:18] It's not that complicated. U.S. can't ignore it. And that's their way to have a part in this ecosystem. So there's a mix of the DeepSeek standard, and then they're kind of like, yeah, this is something that works for us. Let's keep doing it. It's getting them a lot of mindshare and some use in prominent ways. So I think it makes sense.
Inside ATOM: the U.S. response to China’s model surge
Matt Turck [20:34] And is there more of an emerging organized response in the U.S.? I know you're involved, or perhaps behind the ATOM project.
Guest [20:57] I think any concerted response you only see when it actually is public. And I think there's a lot of investment at different stakeholders and conversations that are happening, but that's not that useful. So it's like, I don't have the proof for you, but I do think the right people are talking about it and want to invest more because, realistically, the cost is not that high relative to the trillion-dollar buildout of AI infrastructure. 0.01% gets us better, great open models. Like, we should probably do that.
Guest [21:22] I think that's actually not that complicated. It's just like, how do you get the $100 million line item to the right people that have the talent to do it? And they're like, oh, okay, the right incentives. It's just like, okay, it's a lot. Like, the Reflection AI news is big. It's like, okay, that's probably a good solution for a couple of years. Like, they have enough money and they have a strong base of talent. And it's like, okay, that's like a major checkbox.
Guest [21:35] We need to have diversity there because the Llama thing could happen again where it goes away, but looks like a small snowball, but hopefully grows in the coming months.
Matt Turck [21:46] Today's release and you guys' work is part of that American response to China's rise in open-source AI.
Guest [22:06] I would say I launched ATOM in July and thought it would get more visibility, but now I'm getting a crazy amount of media inbound and press inbound, and everybody wants to share the plight. So it was like, okay, I guess I was just four months too early, but that's what I don't really mean. It's like, just today I saw Bloomberg published a post that pretty much had the same title as my Kimi post from July. I was like, okay, I'm glad that people are paying attention now.
Guest [22:11] It's like, better late than never.
The rise of “thinking models” and inference-time scaling
Matt Turck [22:33] Congratulations. Best form of flattery. All right, switching tacks, and in an effort to make those conversations educational for a broad group of people: one of the key aspects of the release is the thinking model. Could you remind folks what a thinking model actually is versus other forms of models or prior generations of models?
Guest [22:57] A lot of people have heard about inference-time scaling, which makes sense. If you spend more compute at inference time, you get a better answer. A thinking model is really a way to train the model to exploit that a lot. So you spend a lot of tokens, which are usually hidden from the user as a long chain of thought, and the model therefore has this step change where it's way better at math tasks, coding tasks, agentic tasks. I think our future plans are adding more tool use to the models.
Guest [23:26] We're not talking a lot about agentic search or agentic code execution on the fly and stuff for this model. But building thinking models is the gateway to doing a lot more interesting things, like Claude Code. Maybe we'll have OLMo Code next year and all these things that we want to do. The thinking model has just been the thing in 2025 that uses a lot more compute per answer. The model gets way better. I don't like thinking models, but it's fine. No, they're good.
Guest [23:28] They're very useful.
Matt Turck [23:57] Thinking models are really like work mode, and regular instruct models are usually more fun to build. They can be more quirky. But yeah, I think they're like 90% of the cases, especially user-facing cases. Folks are okay spending time waiting for this model to craft a better answer. There's still a space for models that can respond faster. You see stats that Google released about adoption of Gemini Flash, and that's where non-thinking models that can at least approximate, have a good approximate first answer, are really useful.
Matt Turck [24:42] They're also more fun to build. But yeah, thinking models are where the future is, especially when it comes to agents integration. Before we go into the pipeline very specifically of the OLMo family, because, as you alluded to, that's one of the amazing things about open source, is that we can, in a discussion like this, truly understand how the model works versus other conversations with commercial players. So before we go into the pipeline, I'd love to talk a little bit about you guys, your backgrounds, and AI2, which is a very important player in the ecosystem that people may or may not have heard about.
Matt Turck [25:28] So who wants to go first? I sort of stumbled into this role by just picking problems that are interesting. So my background: originally from Italy, moved to the U.S. for a PhD. My PhD is in information retrieval. How do you build a search engine, to simplify it a lot? I slowly got into more and more natural language. After grad school, I joined Amazon. I was working on Alexa, at the beginning working on the search part of Alexa.
Matt Turck [26:01] And then I got, wait, the actual part where the users talk to Alexa? The interesting part. So, slowly moving towards that. Initially, I joined AI2 working on a project called Semantic Scholar. It's still active. It's a search engine for academic papers. And there, the interesting bits were actually interacting with users and less so the actual text of the papers that you were searching on. And then the way I got into LLMs and building language models is really intertwined with how AI2 got into building language models.
Matt Turck [26:42] It all started around, what is it, November of 2022. This is around the same time ChatGPT got released. A bunch of researchers at AI2—this is individual contributors, it was not a direction from the top—a very grassroots initiative at AI2. A bunch of researchers got really interested in building a model that would be fully open. AI2 had already built sort of proto-language models around 2017, 2018. So a lot of the interest was in recapturing, expanding that line of work.
Matt Turck [27:18] So a bunch of us got together, sort of started planning, got in touch with a few companies who might be interested in supporting these initiatives. We got initial grants. AMD, at the time, there was about 2 million GPU hours. And so we sort of had the idea, had the researchers interested, we had the compute. So we went to leadership at the time and sort of told them, hey, we're going to go do this thing. I hope you're okay with it.
Matt Turck [27:50] And one of the nice things about AI2 is, at heart, we are a research lab. So everyone was like, sure, you figure everything out. Just have fun. Great. All right, Nathan, how about you? So you're a man of many talents. You do AI research, you write this very interesting blog/newsletter called Interconnects, you do podcasts, you do a bunch of different things. So tell us about your journey.
Guest [28:12] Yeah, I say I wear many hats to try to get the things that I want to do done. I showed up to Berkeley as an EE PhD admit in 2017, and then I saw that AI was happening and I decided that I wanted to try to do this, which started by going to all the names that people know, like Sergey Levine and Pieter Abbeel, and asking to be in their group. And then they respectfully said no. And then starts the long process of learning how to actually do it without being directly embedded in these elite groups, which was a mix of robotics and reinforcement learning and finding my way there.
Guest [28:45] My PhD was mostly in model-based reinforcement learning, and then my one research job was to go join Hugging Face when they said they were going to make an open-source version of DeepMind to do a bunch of research. Realistically, my job was not that impactful or useful at Hugging Face until ChatGPT came out, and then I was like, oh, I should maybe just learn about RLHF. That got very immediate traction as somebody trying to work in public with the team there.
Guest [29:06] So Lewis Tunstall and other people at Hugging Face are still doing a great job on this, and we worked together for a while. And then mostly I was just getting burnt out on remote work and met Luca in Hawaii at a fun conference and was like, wow, I could have real-life friends. And I joined AI2 to work in person and tried to do the same thing, which kind of takes an evolution of the OLMo story, which was just like I had a lot of motivation in trying to figure out these—what was mostly reinforcement learning from human feedback at the time.
Guest [29:45] And make versions of these post-training techniques public. Then that evolved through both OLMo, and we have our post-training methods named Tulu, which is like, we spent a long time trying to replicate what we thought was close to Llama 3 post-training with multiple stages and optimizers, which is the project that came up with the name Reinforcement Learning with Verifiable Rewards with a bunch of people. So it's kind of this evolving journey at AI2 in search of impact, which is what we think people are actually doing.
Guest [30:18] And then largely, the opportunity that Luca and I and others at AI2 fill is that there's so much money in AI, and it only becomes increasingly so, that the amount of people that can talk about these things in public and educate and get more people involved by spreading knowledge is ever smaller. So I describe my career journey as a lot of it is filling that vacuum and thinking about what's impactful. So it kind of pulls you when there's such a void. It has a sort of gravity to make it clear what you should be doing.
Matt Turck [30:44] You anticipated my question, which is sort of obvious. In a world where we see hundreds-of-million-dollar, billion-dollar packages offered by some commercial AI labs for people just like you, I was curious about your interesting motivation to join AI2, which is a nonprofit. But impact is the short answer, correct?
Guest [31:08] Yeah. I mean, I've been here for two years, and I wasn't famous when I joined. So let that be told to people looking for new jobs: you want to find a job that you can grow into. And I think AI2 has been a really, really good place for that for many people because you have independence and are encouraged to go forth and do things and not be a cog in a broader grind-out-language-model machine, which is important, but it's harder to get visibility.
Matt Turck [31:43] So we alluded to some of it, but maybe a few words about AI2. So AI2 was started by Paul Allen, right? AI2 stands for Allen Institute for Artificial Intelligence. You mentioned some grants, Luca, and I think earlier in the conversation we talked about a recent grant as well. I saw that it was $152 million from NSF and NVIDIA. So what is AI2? How did it start? Who founded it at a high level? AI2 was founded around 2014 by the late Paul Allen.
Matt Turck [32:20] Initial AI2 was very focused on building machines that can do science, can understand science, solve science problems. That's when Semantic Scholar started as a repository of science papers. Slowly, one of the initiatives that started forming was more fundamental research around how language models work, how, at the time, what was called natural language processing was working. You had teams like AllenNLP doing great work. From very early on, we always had this idea of not just releasing artifacts or research, but releasing the tool.
Matt Turck [32:41] Back in the day, we had this very widely used library called AllenNLP that would allow you to build and customize these models.
Guest [33:00] I'm going to jump in. It's cool because it's the namesake of our team name and has been for a long time at AI2. And it was the thing, it was the main competitor to Hugging Face Transformers, and they ultimately outcompeted AI2 as the thing that people use for that because they had a very different model and amount of support. Luca can keep going. Luca knows a lot more.
Matt Turck [33:31] But we have been at it, open-sourcing for a while. I think it's something that folks here understood really early—this is before my time—that it was important both in pushing science and also unlocking commercial use cases that, as a nonprofit, maybe we didn't anticipate. You release a tool, people pick it up and do amazing things with it. Yeah. And we moved on in language modeling more and more recently.
Matt Turck [34:07] Right now, AI2 has maybe like three main projects. One of them is the OLMo model family. And there are variants of OLMo. There are some that focus on the full pipeline, some that are robotics, some that focus more on processing images and video and audio. Is Molmo part of one of those variants? Yeah, you have Molmo. It's one of our projects that work on multimodal inputs. Recently, we released another one called MolmoAct that's more focused towards robotics, receives multimodal input, and then can act in space.
Matt Turck [34:49] And then we had the model that was able to do automatic speech recognition. Another OLMo model focused more on document processing could do OCR. So it's a nifty little family of models. We have a working group on agents for scientific tasks, harking back to our roots. This is the Asta family of initiatives. This is agents to help scientists do their work. And that just came out, right? Like August of this year? Yep. The team has been cooking since the middle of last year, but finally we had our first release this year.
Matt Turck [35:30] There's actually two releases. There was the main Asta release, and then recently we announced a partnership with CAIA, the Cancer AI Alliance, using some of the components in Asta to help researchers make progress on cancer research. And then there is a third branch on AI for the environment, building models that can see and understand so that they can model Earth and can work with different signals to do prediction around the environment and so on. I'm being a little bit vague on this one because I don't know if it has been announced yet.
The full Olmo pipeline, explained simply
Matt Turck [36:06] It's a preview right here. The MAD Podcast is making news. Okay, very cool. Oussama, that's great background. So we've got OLMo, we've got Tulu, we've got Asta. Just maybe one last question: in terms of size, what are we talking about? How many of you guys are there? 200 people between the research staff, engineering, comms, and other support roles. That's fantastic background. Thank you very much. All right, as previewed, let's switch tacks and go into OLMo 3, OLMo Thinking, OLMo reasoning, whatever you guys end up calling it.
Matt Turck [36:44] And I think it's a perfect opportunity to talk about how those reasoning models actually work. In prior episodes of this podcast, we've had great conversations with folks at Anthropic or OpenAI, but not surprisingly, there's only so much they can talk about. And the beauty of what you guys do, which is the very essence of it all, is to make it open and accessible to everyone. So I'd love for us to talk about the whole pipeline, from pre-training to post-training, the different parts, and make that super educational and explain in plain English what part does what.
Matt Turck [37:10] So could either of you start with just a high-level architecture of what the various subcategories of the pipeline are? And then we'll go into those one by one.
Guest [37:32] Sure. I recently gave a talk on this at the Conference on Language Modeling, so I have them on the top of my head. I think I could provide a personal motivation for this, which I think, as researchers, we're closely embedded in the community, and we see that there are a lot of people that are starting to do this reinforcement learning research after DeepSeek-R1 and Qwen3, between 1 and 8B parameters. I think something that particularly motivated a lot of the fine-grained details that we might not have time for in this podcast is that there's some questions hanging over the data used for Qwen when doing this RL research.
Guest [37:59] Specifically, there's two papers. One is "Spurious Rewards: Rethinking Training Signals in RLVR," which is one that I was on with a lot of people at UW and Ai2. And then another one, which was "Reasoning or Memorization? Unreliable Results of Reinforcement Learning Due to Data Contamination."
Matt Turck [38:08] Yeah, actually, let's spend a couple of minutes on that. What do spurious rewards mean?
Guest [38:36] I think the thing to know about this is that you can trigger—you'll hear this rant on the technical side later. It's a lot of background on understanding what these algorithms are. But essentially, the question mark is: did Qwen include training data that is too close to the evaluation targets so that the research is picking up on weird behaviors within the model rather than the fundamentals of what this reinforcement learning is doing?
Matt Turck [38:43] In other words, did they teach to the test versus enabling true thinking?
Guest [39:06] It's a gray zone. I think all the frontier labs will do this to some extent, which is how they're tasked. You have a team member that's tasked with improving an evaluation, and then the easiest way to do this is to train on test. But they all have dignity as elite scientists, where they won't do this. The next closest thing is you do some sort of paraphrasing of the test set to create new training data. Therefore, you're not technically cheating, but you're potentially cheating.
Guest [39:21] It's like, where in the spectrum of you scrape GitHub for math problems versus you paraphrase the evaluation set? Where do you draw the line on actually calling it cheating? Different people have different answers, but mostly I think a goal that we kept coming back to, because we understand that OLMo is not—you can look at the numbers, we're getting close to Qwen3 with reasoning or without—but this is not a 600-billion-parameter model that people are going to immediately download and run OLMo code on or anything.
Guest [40:03] But we want to make sure that our core audience could do the research that we want to do with confidence and debate it. We want to give people access to every stage, and you can then see how this impacts this new important area of research. We're going to talk about six stages. One is large-scale pre-training, which is this training on all of the internet, predicting next tokens. Two is what we call mid-training, which it's debatable whether or not it actually should exist.
Guest [40:35] What technically it is, is you train on higher-quality web data and with a change in the learning rate. Three is long-context extension, which is absolutely essential for these reasoning models because they generate so many intermediate tokens before sharing an answer with you. Luca has a lot of battle stories from that. Then we go into post-training, which in our case—those three building blocks of pre-training are, I would say, more set and super essential. Then post-training, when you approach this, you have a bag of tools, which are optimizers, and you apply them in the order that suits your model depending on size and capabilities you want.
Guest [41:07] We'll talk about things that we did, which is instruction tuning, preference tuning, and then we did some reinforcement learning with verifiable rewards again. But if we were to train a model that was 10 times as big, all this post-training stuff would change. But the pre-training and mid-training and long context would actually become looking pretty similar. It's kind of a difference across two phases of training, where post-training is a bit of an art and you have to do what is best for your specific use case, and that'll change.
Guest [41:18] But we can go through these two.
Matt Turck [41:44] Okay, great. And Luca, you're the pre-training guy and Nathan, you're the post-training guy, right? Is that fair? One of many. One of many. But for purposes of this conversation, and before we dive into each step, this idea of pre-training plus RL seems to be the key idea in terms of progress in the last year or so. And I know the concept of it came up much before that, but in terms of implementation of it, what's the right way to think about it for somebody that's trying to learn about the space?
Matt Turck [42:24] Is one part better than the other, or do they need to exist together? Is RL currently delivering more gain than pre-training? What's the overall kind of high-level take? I think the way I like to think of it is the pre-training phase: it's really like a very expensive initialization of the model. When I think of, oh, what is a good final set of weights that I can pass to Nathan and the rest of the post-training team, it's, well, I want a model that has great knowledge about the world.
Matt Turck [43:11] And it also can sort of—you can start seeing sparks of capabilities that you will want a model that you then want to chat with to have great capabilities. So it is a very expensive and very compute-intensive way to create initial models out of what is essentially random parameters. But it's all about, yeah, let's have this model have a lot of knowledge of world facts and information. And let's have it so that it can start behaving a little bit like a chat model so that when we pass it to post-training and you have this reinforcement learning, there is some behavior to reinforce and to give rewards on so the model can pick it up.
Guest [43:54] I would say that generally the reason why discussions are hard right now on whether or not people should care about pre-training or post-training is that we optimized pre-training for multiple years, and then there was a lot of untapped potential on this type of RL. What is said is that OpenAI figured out a whole bunch of tricks to get o1 to work, and then it showed that this area was possible. And then this year has been a race to capture low-hanging fruit on RL.
Guest [44:21] I think that's kind of the biggest story, is why we have all these crazy new models that appear, like o3, with this thinking and tool use, which are just downstream of, oh, we could do very different things because we have such a good platform of these malleable pre-trained models that we've been iterating on for a long time, that this RL stuff, we just kind of could have tapped into it much earlier. But there's a lot of potential. Yes, the rate of improvement right now in RL is higher, but at the end of the day, it's going to be a dance between both of them, where you need a better base model.
Guest [44:44] It's said very commonly that a better base model and a bigger base model is much easier to improve with RL. So if you take that as one of the core things of doing RL research, it's pretty obvious that pre-training is very important to enabling that.
Matt Turck [45:10] As a quick detour, there's been that podcast with Richard Sutton that was effectively saying that RL was the way to go, and that pre-training and LLMs was a little bit of a flawed premise because it was sort of an imitation of reality, basically doing the way humans described reality as opposed to being confronted with the actual reality through RL. Do you guys have any quick take on that?
Guest [45:35] My take is that a lot of people are being exposed to Rich Sutton for the first time, and Rich is a font of wonderful ideas, but often not ones that are going to be immediately practical. This is how you get things like creating reinforcement learning, but not necessarily things that are going to impact what GPT-6 is. So I've been on the critiquing Rich train for many years before this in terms of making people try to interpret his ideas as realistic. I think the one from 2021 or 2022 was his Reward Is Enough paper, which essentially is an argument that a reward function is sufficient to get any intelligent agent that you want.
Guest [45:54] So I think that that's actually, rather than the technical debate, is an entertainment of the whole community being nerd-sniped for the first time by that.
Matt Turck [45:58] Distraction.
Guest [45:58] Okay.
Matt Turck [46:32] The message is not that surprising. There's this fine line between the actual ideas, and then there is the engineering around it. A lot of making language models work is engineering, and not in a denigratory way, but in a way that's like, we got to figure out how to translate a research idea into practical things. And so pre-training is just a good way to initialize one of these models. If better ideas come out in the future, we can switch to that. No one is married to LLMs being the end-all solution.
Pre-training: data, scale, and avoiding catastrophic spikes
Matt Turck [46:57] There's a big difference between just describing the system in theory and then actually getting them to work. Thanks for that. So let's take those six modules turn by turn. So let's talk about pre-training. What did you guys do specifically for this model? Pre-training is very interesting. The way we sort of plan—so, a good background to have is that pre-training, all that happens during pre-training, we have to be very methodical in how we do it because, first of all, it takes a long time to pre-train.
Matt Turck [47:39] I think it's standard practice among the frontier labs to try to cap your big final pre-training run to two months, not more than that. But to get to something that will not crash and burn during these two months, you have to do a lot of preparation around this. So really, everyone who works on pre-training is fairly methodical.
Guest [47:47] Okay.
Matt Turck [48:13] And just to sketch out how that works, it's usually you have a sense of, okay, the duration of this run is fixed. The number of GPUs I have available will be fixed. And therefore, you write the fastest possible code to train this model. With these three, you can figure out, okay, how much data can I show my model? In our case, that number was like six trillion tokens.
Guest [48:14] Wow.
Matt Turck [48:36] Given that number, then we go back and we figure out, okay, what are the best six trillion tokens out there? And the way you figure that out is a combination of what data you have access to. We want to eventually release the data, so we limit ourselves to data that is publicly available. So either internet text or PDF documents that you can find on the internet, or code that you can find on the internet.
Matt Turck [49:11] And then among this pool, our initial pool was closer to 300 trillion tokens. You shrink it down till you reach your target number, and hopefully, as you shrink, you only keep the best part of this. So you remove duplicates. You have ways to judge, is this document better than this other document? We have a way to evaluate the capability of the model. So if your evaluations want medical documents because there's a medical task there, you figure out how do you pick documents that have good medical information.
Matt Turck [49:45] It may be at the expense of some other domains. But yeah, it's this delicate balancing act to find this data. And after you commit to this initial run, you will do your training of this run. Over there, there's a lot of making sure that the way you design the model doesn't suddenly start forgetting what it's learned. We call these spikes in the language model, but basically, you don't want this event that, if it happens, you have to restart from scratch and you can't recover.
Matt Turck [50:24] So there's a lot of work on that, but after these months of training, you get to a final model, and then on this final model, it will still lack some capabilities that I know Nathan's team cares about. So these are things like long context or being able to solve some problems to start with. That's where things like long-context extension or mid-training happen. Yeah, let's get into that. So that phase two, so mid-training. So again, a term I personally hadn't heard of before.
Mid-training (tail patching) and avoiding test leakage
Matt Turck [50:52] And Nathan, briefly describe what it is. Double-click on that. I heard that some labs, instead of mid-training, call it tail patching, which I think is a much better term. And the term is, like, at the tail of training, at the tail of pre-training, you patch the model so that the things it hasn't learned in pre-training, you learn after. You learn at that phase. And, of course, when you do that, you also need to make sure that the model doesn't forget stuff it's seen during pre-training.
Matt Turck [51:25] So that's why you mix in some of the best data from pre-training that you carry over. So you give it more code data, for example, or math data, that kind of stuff. If the model maybe cannot reason about certain math problems, you do it. That's like when Nathan mentioned earlier, sometimes there is some leakage of things that look like the test during this phase. There is an uncharitable way to describe it, which is like, oh, someone is trying to cheat there by adding this data.
Matt Turck [52:01] But it's also, like, it's so easy to accidentally leak your test data in there. We spend a lot of time making sure that doesn't happen because it's really a tricky balance, because you want the model to start being able to solve problems like the ones that you see during tests, but you really don't want that test data to accidentally leak there. Otherwise, you can't measure how well your model does. And then you mentioned long context, which is the third stage in the six-stage pipeline.
Why long-context training matters
Matt Turck [52:35] So why the focus on long context? And I guess, what does long context mean in the first place? You want these models to be able to work with very long sequences of text, both as input. Imagine you want to give it, I don't know, a collection of documents. And you also want this model to be able to generate a lot of text in the output, especially now that you have these reasoning traces, right? These thinking tokens. Why don't we train from the beginning the model to be able to do that?
Matt Turck [53:08] It's because the longer the input that a model is trained on, the slower it is. The rate at which it gets slower is higher than the length of context. It's a quadratic slowdown. So we definitely don't want to do the entire pre-training at this extremely long sequence. But at some point, we have to teach the model to actually work with these long sequences, and we save it for the very end so that we can do it in an efficient way. And I think you mentioned somewhere that data doesn't matter for long context.
Matt Turck [53:20] What do you mean by that? And then what does matter? This is getting a little bit in the weeds.
Guest [53:30] No, Luca loves data. Luca likes to be in a dark room grinding out tokens to train the models. That's the emotional backdrop for this. It's very painful.
Matt Turck [53:54] It's very technical stuff. Do I use QK-Norm? Do I use GQA? It doesn't really matter. But there are technical decisions in how you set up your model that—you can have the best data in the world, and your model will not be able to reason over many, many tokens. So it doesn't matter in the sense that you can train the model on bad data. You can have the best data in the world, but if you set up your model wrong, you're never going to recover it.
Matt Turck [54:32] So sadly, I can't be the savior with the magic tokens that makes the model good. We have to make the model with the right architecture. Okay, so that's stage three, long context. Maybe just to bring this to life, what's the difference between before and after? If you have a 40-page PDF that you feed into the window, will it just get faster results or better results? What happens? At the beginning, you just can't do it. You pre-train at something like 4,000 to 8,000 tokens.
Matt Turck [55:06] That's what we use for OLMo 3. That's what LLaMA used. That's about maybe eight pages if you use double spacing, newline kind of thing. And after that, we extend to about 65,000. In industry, you have extensions of a million tokens. I think Gemini recently announced, like, over a million tokens. At that point, a million tokens is like 10 books. So you can work with an extremely long amount of information. It's nice. You don't have to think about, if you're building an application with this language model, you don't have to think, like, oh, of this amount of information, how the heck am I going to extract the ones that I need to show the model?
SFT: building the foundation for reasoning
Matt Turck [55:47] Can just give it all, and the model will figure it out. So it really unlocks a lot of opportunities. All right, so that's the pre-training world between pre-training itself, mid-training, and long context. So now let's switch to the post-training world, Nathan, if you will. So starting with SFT, which stands for supervised fine-tuning.
Guest [56:11] Yeah. One of the things, especially for a model like OLMo where we're scrappy and putting everything together over time, is that one of the biggest changes is that when reasoning models become popular, the in-vogue evaluation suite of the industry shifts to add a whole bunch more new things in. So one of the things that happens at every stage is that even if a lot of the data has overlap, you mix it in a different way. I think Kyle Lo and Mei Chen, that's another researcher and an intern, did this whole mixing procedure that we use across all these stages.
Guest [56:41] Just to upweight the math, code, and reasoning stuff to make sure that what happens later in post-training is much more tractable and that all this stuff is set up. That's the type of thing that we have to do that's baked into everything. Then post-training, I think for this model, everything we're doing is operating on the assumption that this is about a 7B model. We are very narrowly focused, and therefore we're going to do what many people have done, which is called distilling from bigger teacher reasoning models.
Guest [57:16] I think distillation is described as when you take the outputs from one model and then you fine-tune on it later. I think there's been a lot of broader discussions on this in the community. Then this supervised fine-tuning stage, or SFT, or instruction tuning, is all about just getting the best traces from reasoning models out there or the community and then just teaching your recently trained base model to behave really, really closely to what is going on there. In our case, we took a mix of existing datasets like OpenThoughts-3 and modified it, which is from Bespoke AI Labs, a startup.
Guest [57:37] Then we also generated a whole bunch of new data. So we ended up using a mix of teachers from DeepSeek-R1, DeepSeek-R1-0528, which was their updated version, and then Qwen's reasoning model, QwQ. These tend to be pretty strong teachers.
Matt Turck [57:50] Why is that? So you have a pre-trained model, but for supervised fine-tuning, you're still using a different model. Why is that, in simple terms?
Guest [58:13] Essentially because our small model is not going to be able to output as strong of text. There's a kind of fork in the process where I'm talking about a small model. If we had a bigger model, what we would do is do a lot of reinforcement learning to start, and the model then would take time to learn these interesting behaviors and have strong performance. But with a smaller model, the ceiling on that is fairly low. It just doesn't have the capacity to learn from these harder math problems. So what the common practice is, is you take the absolute best reasoning models you can get that are openly available with a good license, where you can just generate new data yourself and train on it and release it to the community, which is something we've been seeing a lot of this year.
Guest [58:42] Therefore, the models that are closest to the frontier in performance with a good license all happen to be Chinese models throughout the year for this case. I think in our case, even if GPT-OSS had existed, I don't think we would have used it for synthetic data in this because that model is really designed for tool use, which is something that we did a bit of in this project, but not in the sense that that model is, which is this many-hop agentic reasoning with search and stuff.
Guest [59:18] The DeepSeeks and Qwens of the world are just powerhouses at generating math and code answers and other things and being generally robust. Five million reasoning traces, mostly on math and code and STEM, but also on chat and other general capabilities. The model really absorbs a lot at this point. I think if you were to have told us last year, when working on OLMo 2, looking at it, that if we had a similarly sized OLMo model that gets like 95 on math and like 70 on AIME on these crazy math evals, it would have been surprising.
Guest [59:54] But this is just what you can get when you can extract data directly from these really powerful models and distill it down. I think realistically a lot of companies are going to want to do this because you can do this for your domain. I think we threw a blanket on, we want all of these evals from instruction following and make sure that you can actually talk to the model and not have it just become totally broken. But you can do this in any specific task you want if DeepSeek has coverage on it, and it's very efficient.
Matt Turck [1:00:04] Okay, great.
Guest [1:00:30] That is the foundation. If you're training a small reasoning model, you need to do this. Then the other things after are how do you extract more performance, and they quickly become more technical or done because we want to do them and maybe not efficient in our time. So 90% of our time this summer is having great people battle reinforcement learning infrastructure because when you generate a lot of tokens, the time or compute increase and memory increase is quadratic. Therefore, you pretty much encounter every possible bug in your framework or every possible corner case that'll make your job go to a halt.
Guest [1:00:59] But most of the performance is through this SFT and the preference tuning that comes before it. But the RL is like, we need to do this in order to build the infrastructure for many of the future OLMos that we want to build later this year, where they get bigger and they can do more interesting reasoning with tools and so on. So it's kind of like a nuanced point of the model. It's like, yeah, we'll show you that we got a couple of points out of doing RL at the end, but really the RL tooling is something that is so crucial to doing the next models that come from here.
Matt Turck [1:01:37] Right, right. And thank you for that. And just to, again, in an effort to make this interesting to just a broad group of people who are curious to understand how AI works: so SFT is not RL yet, right? That's supervised fine-tuning. So that means that you basically show the model a golden copy of what good looks like, and you train it based on that labeled data. Is that a good way to describe it? Yeah.
Guest [1:01:55] So it's the same loss function as pre-training, which is you're predicting the next token. In this case, what it looks like is a question could be like an AIME-style, really hard math question. It would be like, list all the prime numbers within some constraint of x and k. And it's like this one sentence that is really hard, and then the model generates 30,000 tokens of, let me think about this and do this, and to test this, I'll have to use this theorem and hypothesis.
Guest [1:02:35] We were talking about token intuitions for a bit, but 30,000 tokens to solve a math problem is pretty mind-bending. So if I were to sit there and read this, it would be hours of me just trying to read this one math solution. So these models are very unintelligible in many ways. I think the reasoning models sometimes will go into a bout of guess-and-check for hundreds of attempts before realizing that they can no longer guess-and-check. I mean, this is our reasoning model.
Guest [1:03:08] I think the frontier models could have probably done this and fixed this issue, but there's just really, really, really odd things in these tokens. But even with that, doing this next-token prediction is an incredible foundation of performance that many people use. So it's not matching any sort of human reasoning or things that people might want it to be doing, but it is teaching the models their own language of breaking down problems step by step in order to solve a goal.
Matt Turck [1:03:22] For this specific stage of SFT, do you want to talk about how you went about creating the dataset for it? So, precisely this representation of what good looks like for the model.
Guest [1:03:24] Yeah. Luca, do you want to jump in? Do you have things too?
Matt Turck [1:03:53] The other thing I was going to mention is that sometimes in the big announcements of the frontier labs, you don't see what Nathan was describing around having to do SFT to then do RL. It's very common. We're in a common situation where this is uncharted technology, right? So you have nothing. You have to find ways to fix some components of your pipeline before you can build the rest of your pipeline, and then go back fixing the first part. So for us, it's, okay, we want to do reinforcement learning on this larger model.
Matt Turck [1:04:22] Okay, we need our reinforcement learning code to actually be super fast, super reliable, and useful. If we need to iterate on that part, we want to iterate with smaller models first because we can iterate faster. They take less compute to work, so we can do more things in parallel. Smaller models, they cannot do RL first. You've got to create the data first. You have to go to the SFT. And then we are lucky enough that there are other great models that are open source that we can use to create this data, versus the alternative would be, I don't know, to spend—it's not even the money—to spend a lot of time instructing humans to create the same volume of data.
Preference tuning & why DPO still works
Matt Turck [1:05:02] Slow things down. So it's a lot of this of, I don't know, you're building the tracks as the train is going down at incredible speeds, and you have to figure out ways to fix some parts of your pipeline so you can work on the rest. All right, so let's talk about the next stage in the pipeline, stage five: DPO and preference tuning. What is that? What does that do?
Guest [1:05:25] Yeah, so this is one of the things that is thought of as like, hey, let's try this. We're not sure if it'll work, kind of later in the process, when you spend a lot of time on other things, and it works very well. I think DPO, or direct preference optimization, is not exactly new. I think it's a way of optimizing for preferences. It's related to this whole RLHF thing that we mentioned. Technically speaking, in one sentence, it's an analytically derived loss function that is essentially applying stochastic gradient descent to the RLHF objective.
Guest [1:05:57] So it becomes much easier to implement than other things. We used this in the past with OLMo 2, with Tülu 3, Tülu 2, other OLMos. The question was, can we apply this out of the box on top of a reasoning model? We knew that it works in many different situations, because we weren't sure what would happen with these long reasoning traces being included in the loss function and so on. So then, essentially, there's a student, Scott, that has been working on what he calls the delta learning hypothesis, which is an intuition for understanding DPO as being more about the contrast between your chosen and rejected examples.
Guest [1:06:28] The core of preference learning is that you have pairs, or some grouping, of completions to the same prompt. You have one question with multiple completions, and his intuition and work is showing that this contrast is more important than the exact magnitude of goodness of the answer. So what he did is he spent a lot of time trying to come up with a good pairing of reasoning models, which are open source, or open weights, and they have a permissive license and they include the reasoning traces, because we kind of need this, and you need them to be sufficiently well spread out.
Guest [1:07:11] So we spent a bunch of time generating this data and doing some normal kind of like, let's fiddle with the learning rate and small things. And it was kind of just like, yes, this works. After we did it, we saw that Hugging Face did something similar with SmolLM. They trained a fully open 3B model where they pre-trained it as well. The funny thing is that we converged on using the same Qwen 32B and Qwen 0.6B. The problem is that these small Qwen models and these small public reasoning models are actually so strong that getting a sufficient delta to another model to apply this preference learning technique was kind of hard.
Guest [1:07:53] Our past techniques, we kind of had groups of models we sampled from, but as these open models are getting better, these samples become too homogeneous for the learning signal to exist. So it's a cool experiment because it validates this hypothesis of the changing tides. If you think about years ago with Alpaca and stuff, those models were so broken that having this group had enough variance and contrast in it where we could do a different type of preference learning. Where now, you have to look really closely at the completions and make sure that there's a learning signal for the models.
Guest [1:08:19] We did this and it kind of gave us a boost across the board, I think. Sometimes things look very easy when you've done careful data work and set up to understanding your optimizers. I think Luca described pre-training as very scientific and post-training as the Wild West. I think there's many analogies. So it's like, we had me that made this SFT dataset where I was like, we had a bunch of cloud credits and they were running out and we were behind, and I just generated as many completions as possible.
Guest [1:08:53] So it's a few billion completions from DeepSeek over the weekend. You're like, oh, we'll mix it and filter it later. I applied filtering and the answer was like, oh, we just include almost all of it. We did very little. We would have liked to do more if we had resources for longer, but sometimes there's low-hanging fruit and doing the obvious thing yields a lot of results. This SFT and DPO stage, in a lot of sense, are that, and then this RL stage is extremely hard technical grinding, week in and week out, to make the tools even run at all.
Guest [1:09:30] And the disparity in post-training is like, yeah, that kind of tracks to me. You just have all these checkpoints flying around and it seems like chaos, and then something that's extremely obvious gives you a massive gain. GPT-5 level to almost Qwen 3 level. It's just the thing that you apply to get there is sometimes really obvious. And I think the frontier labs are much further down this path, where they take these low-hanging fruits so fast. But as a smaller team that's trying to map to what the changing priorities of the field is, sometimes it's just turning the crank on this really straightforward thing.
Matt Turck [1:10:10] I don't remember if I said it, but Dario from Anthropic said very plainly, look, what works here with 50 to 100 lines of code. He was saying it in the context of espionage and him being scared about some trade secret from Anthropic being exported out. But the solutions, at the end, everyone in the industry favors are actually very simple. The problem is that there is a very large space of equally simple solutions, and all the work goes in like, okay, how do you test these?
Matt Turck [1:10:46] How do you test them as fast as possible? How do you convince yourself that these results look good? They're not just because, oh, sometimes there's a bug somewhere that causes something to be too good to be true. So yeah, a lot of it is less about the final solution, what matters. It's about the speed at which you iterate and how robust your tools are. So immediately after you see good results, I know that this is a good result.
The hard part: RLVR, long reasoning chains, and infrastructure pain
Guest [1:10:52] All right.
Matt Turck [1:11:18] The journey since we started talking about RL. So, RLVR, reinforcement learning with verifiable rewards. Let's spend a little bit of time on that sixth stage in particular. Nathan, I understand that's your baby, or you're one of the fathers of the baby. Do you want to walk us maybe a little bit through the history quickly?
Guest [1:11:47] I mean, I think that I'm the person that got to bring it publicly to the world. It's well known that people across industry have been doing this for years. And then the technique started to get far more impactful. It's broadly taking existing reinforcement learning algorithms, or downstream evolution of proximal policy optimization, PPO, which is an evolution of REINFORCE. And then DeepSeek had their group relative policy optimization. I always try to say group robust. I think it's group relative. And all these algorithms are really quite similar, and you're training the models with whether or not they got the answers right or, in the case of code, whether or not the tests execute and don't fail.
Guest [1:12:25] I think one of the famous examples is that doing too much of this, or racing to get the low-hanging fruit from this RL approach, is what makes all these code models do all these try-except things to avoid errors because they accept all the errors. I think that is just because the gains that you get in the model being useful are so much higher than the annoyance and the fact that it also does these stupid things. And we'll fix the stupid things eventually. In the case of this OLMo model, it's not anything crazy.
Guest [1:12:49] We cast a wide net on RL math problems. We do some data comparisons to see which data we think is the best for teaching these models. We do mixing with code and precise instruction following. This mixing is effectively when you tune the big set that you have to what you've known from many experiments and to the specific model checkpoint that you're on. So if you have a really strong model and you show it really easy math problems, there's no learning signal.
Guest [1:13:08] And if you have a really weak model and you show it really hard problems, it gets them all wrong. There's no learning signal. So the learning signal is all from the gradient of: you sometimes get it right and you sometimes get it wrong.
Matt Turck [1:13:14] Do you want to give the plain English definition of RLVR versus RLHF?
Guest [1:13:35] RLVR, the verifiable rewards, is in the name. I think essentially the reward that you get from the environment, which is the completion or the grader, is whether or not you got the problem right. In RLHF, the reward is essentially a reward model which is rating the quality of the response based on a proxy to what humans would like. So it's described as being much—the RLVR reward is much easier to understand because these reward models tend to have a lot of problems, and you can over-optimize them much more easily because the reward models will pick up on features that are maybe emojis or something like this that you don't actually care about, whereas RLVR is much better matched to performance characteristics rather than style.
Why RL is so technically brutal
Matt Turck [1:14:13] You tweeted, I think, or said somewhere that RL and long-context reasoning distills is very hard. I don't know if that's RL in general or specifically this type of RL. What makes it super hard and very much the frontier of AI right now?
Guest [1:14:39] There's many ways that your tooling could fail. I think where most of these processes are set up right now is that you have a set of generation GPUs, which look like something like vLLM, and you have a set of training GPUs, which is some distributed learning framework, which is where you actually have this RL update and loss function. Therefore, you need to have some sort of system that orchestrates the two and passes information back and forth. This kind of information passing back and forth is really annoying.
Guest [1:15:06] It's a systems problem because you have distributed error handling and things like this. A common case is when you have the most basic approach: you'll have one generation, this one math problem, the model is thinking and thinking and thinking and thinking. So you have all these GPUs working on one problem. Effectively, your whole system is somewhat idle, waiting for the answer. There's many other small things like this, which is this long-context generation just uses so much memory that you then need to introduce different types of parallelism and stuff to do the generation effectively.
Guest [1:15:45] And there's just a lot of subtle numerical issues. So I think it's just kind of stress-testing a lot of the post-training infrastructure that we have had by turning up a lot of different things that could go wrong to the maximum. I think the things that the open community struggles with is that vLLM and Hugging Face use different kernels to do the actual internal computation of the model. These kernels are the things that make things like vLLM really fast. But this then results in subtle numerical differences between the completions that you're generating from the model and then the log probs that the thing that's doing the loss function actually generates.
Guest [1:16:22] If you look at the math of these RL algorithms, it's assumed that those are from the same distribution. Therefore, this is a big root cause of a lot of numerical problems. Then if you look at what we're doing, and a lot of labs have done throughout the years, they do different fixes to change these dynamics. Thinking Machines had a famous blog post, one of their first blog posts, on deterministic vLLM to make it exactly deterministic. That is really useful, and people think it is key to their Tinker API and doing other sorts of RL things where you just have complete control over sources of nondeterminism, and that could just be numerical lack of robustness in RL.
Guest [1:17:03] There's also these discussions on if the open labs have worse algorithms than the closed labs. In reality, it seems like most people are using something like an evolved version of GRPO, which is a bit simpler than PPO. Some labs might be using a learned value function. It's not that important about the details, but what happens is that each lab finds the set of tweaks that they need to get really stable RL performance. In the RL literature, historically, there's a pretty low bar on the amount of changes that are needed to call it a new algorithm, but it's realistically an implementation detail.
Guest [1:17:41] Everyone finds their stable configuration for operating, and it's really dependent on the tool. Yes, you could say that they have a different algorithm, but it's also not really something that you could easily exfiltrate from a lab because it's dependent on many layers of the stack and maybe what chips they're operating on. So it's just one of these things where training these models is complex, and the kind of quick quips could never reflect that.
Matt Turck [1:18:13] The stack for post-training is also so new. Software-wise, I feel like the big strides started happening in 2024 around, you do both: you train the model, but also you run the model at the same time, and they have to happen at a certain cadence, versus pre-training, the seeds of distributed pre-training, like you had in TensorFlow, which Google released in 2017. Right. So it's a much more mature stack versus what you need on the post-training side. All right, so maybe as a last part to this conversation, it's been really fascinating and illuminating, everything that you guys have described, because in particular, it sort of highlights the complexity of the systems, like the multiple stages. And I love what you mentioned, Nathan, a few minutes ago when you said the pre-training is scientific and post-training, my words, not yours, but my interpretation of your words was, like, it's a lot of tinkering and putting things together in a way that you hope is going to work and truly diving into how those models work on the one hand.
Complexity tax vs AGI hype
Matt Turck [1:19:24] But on the other hand, each time you open a newspaper online or go on Twitter, everybody's talking about AGI. AGI and how we're almost there and how it's going to change everything. There's a little bit of a cognitive dissonance between the reality of trying to make those models work, with all the unbelievable progress that we've seen, of course, on the one hand, and the discourse on the other hand. Nathan, you've had a much more, I would say, tempered view of AI progress compared to some AI researchers.
Matt Turck [1:19:40] You had a great blog post very recently that you called "Thoughts on the Curve." I'm curious what your latest thinking is. And look, obviously, feel free to jump in any time, but in terms of what you described in that essay and in a prior one as complexity and complexity tax, which, again, in view of the pipeline you just described, one starts to understand the level of sheer complexity. Yeah, so ultimately, I definitely describe myself as lightly AGI-pilled, and I think you have to be to appreciate the magnitude and gravity of the situation that we're in.
Guest [1:20:37] But also, I think that I'm very far from believing in any sort of singularity being possible due to these things like complexity. On one hand, we talked about all these things which are low-hanging fruit to improve the models, and I don't doubt that—I mean, Sholto commented this on the pod and other places—where at these labs they still see low-hanging fruit in improving the models. In many ways, I don't think that their approach feels that different. They've just refined it relative to what we're doing.
Guest [1:21:02] But at the same time, as things get complex with tools and adding more layers to the stack, and you have to build a product to scaffold it, if the requirement to get the best out of Claude is to use Claude Code, which is some magic product in prompting relative to GitHub Copilot, this is one thing that you're going to need to get right in order to get AGI, along with all these tool uses and stuff. So, as any system gets more complex, the pace of change is slower.
Guest [1:21:29] I think any tech company has seen this. And then, realistically, there's going to be physical constraints on the amount of infrastructure that we can build. So this belief is simultaneously giving us these new data centers and, I personally think, hopefully new power generation, but there's a cap. All these things are plateauing, and then you 10x the compute and you get a big jump. You can't do this forever. So, realistically, there's going to be some physical constraint that kicks in at some point.
How everyone can contribute to the future of AI
Guest [1:21:58] But balancing that with complex systems and the low-hanging fruit results in—I think these researchers are going to grind out improvements for multiple years, but never in a way that results in this kind of accelerating well that we get drawn into. So it's like, I don't know, in some ways it feels like I'm having my cake and eating it too, but it seems like the likely outcome if you look at other types of technology.
Matt Turck [1:22:16] Yeah. And that conversation was in particular reaction to AI 2027, which is a really interesting conversation where the—I'll let you summarize the premise, but the short version is that AI sort of builds itself and therefore accelerates.
Guest [1:22:39] Yeah. And I think they have these milestones, like AI automates research engineering, and then AI automates AI research and development, where it's like each of these are incredible jumps in performance. And I think what's more likely is this messy co-evolution. They deserve credit on their marketing and getting this impact, for sure. But even they are now like, oh, maybe we should have called it AI 2028 or AI 2029.
Matt Turck [1:23:11] So I think that is the reflection of there are these real constraints, but the progress is also going to—normally, it is true, the growth in capability of these models, but working on one, I think it's very unlikely that we will see a discontinuity at any point. That has nothing to do with whether we'll get to a definition of AGI or superintelligence that people are happy with. We will get there. It seems unlikely that the moment—it's going to be a looking-back exercise of like, oh, these were the important milestones and this is what really worked.
Matt Turck [1:23:55] Building this model is so much a collection of refinements to unlock the next stage that it's going to be this smooth trajectory. Whether it hits at some point where we don't have more capacity to keep improving or whether it forever accelerates, people are going to be disappointed if they want to see a moment where one day they log into Twitter and AGI is there. It's messy, and it's fun working on it because it's messy and it gives you a lot of satisfaction. So to play it back, you're both saying yes to AGI, but no to discontinuity/singularity.
Matt Turck [1:24:15] And one, is it fair? And two, if that's what you're saying, then for AGI, using the current paradigm, basically what we just described in the last hour of pre-training plus RL gets us there?
Guest [1:24:47] I think the AGI word is actually pretty not useful. I think that how I describe it is that big tech has all collectively realized that these language models plus scaffolding is going to unlock absolutely incredible value. And I have very high probability, barring extreme geopolitical situations, that big tech executes on this vision across the two to five years to build 95 to 98% of the way there of what you can do with our physical power constraints and what an LLM's ability is. And I think that that will be extreme.
Guest [1:25:17] The transformation from that by 2030 is going to be so powerful across society. There's a bunch of long tail. There's going to be mass societal readjustment to what the internet and media and information means within five years. And that's mostly why I do this. And I think debating whether or not it's AGI is kind of secondary to the fact that this is coming and we want people to study and understand what is happening.
Matt Turck [1:25:30] And to that last point, what does that mean, study and prepare? What would you recommend people do? Although if people have made it all this way to this point of the podcast, they've already done a bunch of the work.
Guest [1:26:01] So I think there's a lot of interest in AI outside of the CS majors and AI people of the world, where it's informing policymakers. I think it still takes a long time for information to diffuse, and there's often not that many people that are engaging in this that are doing it just purely for this kind of—you can call it alignment and concern. There's just a lot of general noise, and I worry about concentration of power or all sorts of many things. And it's just trying to upskill people into understanding AI so they can be engaged, engaged citizens and think about how it affects their domain.
Matt Turck [1:26:38] The other part, maybe on a more positive note, is if the scaffolding is what really moves a lot of us from a broad-capability model to something that actually has meaningful impact, that scaffolding is not just like, oh, only the labs of people who train models can do it. The number of people who can contribute to that, both in terms of people with technical expertise and people with non-technical expertise, is much larger. If the scaffolding is what really moves capabilities, what gets us to this incredible technology being realized, then the number of people who can contribute to it is not just those who work at frontier labs.
Matt Turck [1:27:22] There's a tremendous amount of technical work to do, but also non-technical as soon as you start integrating this technology in the life of real people. As soon as you start working on high-stakes medical applications or other high-stakes domains, then a large amount of the population can contribute in making this technology better and make it work for everyone. Just the base model, I feel like the number of people who can really help make this technology really work for everyone is large. Everyone in society feels like they can contribute.
Closing thoughts
Matt Turck [1:27:51] All right, well, that feels like a wonderful place to leave it. Thank you so much, both not just for this conversation, but for all the work that you're doing in open-source frontier AI, which feels sorely needed and extremely important. So, really appreciate it. Really appreciate the time and all the thoughts. Thank you so much. Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching or listening to this episode from.
Matt Turck [1:28:11] This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.