State of LLMs 2026: RLVR, GRPO, Inference Scaling — Sebastian Raschka

The MAD Podcast with Matt Turck · with Sebastian Raschka, AI researcher and author

Sebastian Raschka is the AI researcher and author. We cover why Transformers remain state of the art despite cheaper alternatives, how GRPO makes verifiable-reward training cheaper by removing reward and value models, and why inference-time scaling, tool use, and prompt management can materially improve model performance without changing weights.

Watch on YouTube

Chapters

  1. 1:05 — Are the days of Transformers numbered?
  2. 6:01 — Small “recursive” reasoning models (ARC, iterative refinement)
  3. 9:45 — What is a diffusion model (for text)?
  4. 13:24 — Are we seeing real architecture breakthroughs — or just polishing?
  5. 14:05 — World models: what they are and why people care
  6. 17:26 — “Pre-training isn’t dead… it’s just boring”
  7. 18:03 — 2025’s headline shift: RLVR + GRPO (post-training for reasoning)
  8. 20:58 — Why RLHF is expensive (reward model + value model)
  9. 21:43 — Why GRPO makes RLVR cheaper and more scalable
  10. 24:54 — Process Reward Models (PRMs): why grading the steps is hard
  11. 28:20 — Can RLVR expand beyond math & coding?
  12. 30:27 — Why RL feels “finicky” at scale
  13. 32:34 — The practical “tips & tricks” that make GRPO more stable
  14. 35:29 — The meta-lesson of 2025: progress = lots of small improvements
  15. 38:41 — “Benchmaxxing”: why benchmarks are getting less trustworthy
  16. 43:10 — The other big lever: inference-time scaling
  17. 47:36 — Tool use: reducing hallucinations by calling external tools
  18. 49:57 — The “private data edge” + in-house model training
  19. 55:14 — Continual learning: why it’s hard (and why it’s not 2026)
  20. 59:28 — How Sebastian works: reading, coding, learning “from scratch”
  21. 1:04:55 — LLM burnout + how he uses models (without replacing himself)

Transcript

Are the days of Transformers numbered?

Matt Turck [1:00] Hey, Sebastian, welcome.

Sebastian Raschka [1:05] Thanks for inviting me on your podcast today. I'm excited to talk about anything AI, I guess.

Matt Turck [1:37] Wonderful. So we are going to go into the state of LLMs in 2026 in depth, including very much post-training and reinforcement learning. But I wanted to start the conversation with the Transformer architecture itself, obviously the backbone of the entire generative AI revolution, but also over eight years old at this point. And for all the tremendous progress in LLM-based systems over the last year, it also seems that there have been some interesting developments in terms of alternative architectures. So, has anything caught your eye?

Matt Turck [1:46] And do you think that's a world where the days of the Transformer architecture could finally be numbered?

Sebastian Raschka [2:09] Yeah, that is actually a very interesting question to start with. I mean, starting at the very beginning with the Transformer architecture, you said eight years. I think it's almost nine years. It was 2017, quite a long time. And I think the question you raised is if it's, like, the final architecture. I think people probably ask that every year: is this the thing we should be betting on going forward? Let's say in 2026, I would say right now, yes, because it's still the state of the art.

Sebastian Raschka [2:43] So there is nothing really better in terms of state-of-the-art performance, getting better-quality results. What we have seen so far, though, are alternatives that make it cheaper. So they have tricks to make the architecture cheaper itself, like linear-attention variants that are, like, a building block in the Transformer architecture. A big one was mixture of experts, which is essentially making the model bigger without necessarily making it more expensive to use in inference, like keeping that reasonable while expanding the size. You see all kinds of, I would say, levers, tips and tricks, hacks around that architecture, but it's still kind of like the same architecture at the core.

Sebastian Raschka [3:16] It's not like a big leap. It's still the same scaffold. At the same time, you have other alternatives popping up, like diffusion models, so text diffusion in particular, or Mamba models, state-space models, and so forth. They all try to address a problem that the Transformer has, namely that it is expensive and big and expensive to run and train. But then, of course, there's no free lunch. These have other trade-offs. They are cheaper to run in certain instances, if you take a look at diffusion models or text diffusion models, but then you don't get the same, let's say, quality out of it.

Sebastian Raschka [3:48] And if you want to get the same quality out of it, in this particular case, you have to crank up the denoising steps, and then you end up with something very expensive. So right now, I think we're at that point where there is no free lunch. We are still trying to figure out what is the best next architecture. Right now, there's nothing on the horizon that would replace that. So my short answer is, I would say right now, if I were to build a state-of-the-art model, it would still be a Transformer-based model.

Matt Turck [4:08] Great. What do you make of world models?

Sebastian Raschka [4:24] Yeah, world models are also an interesting hot topic. So there is the whole world-model aspect for more, like, images and physics and that stuff. So world models are basically models that have, like, an internal model of the world. So they kind of simulate something internally, what you have externally. Like, for example, if you have, like, a chess-playing model, it has, like, an internal chess simulator built inside, so it can kind of make better predictions or predict the next states.

Sebastian Raschka [4:55] I think that's particularly interesting for robotics. But coming also back to LLMs, there was also a paper by Meta. It looked very promising to me as a refinement or a next step for code-based LLMs. LLMs for coding are still next-token predictors. But in addition to that, what they did is they also tried to predict the internal states of the variables during training, an objective to, if you have Python code, say, okay, at this iteration, if someone would step through the code, this variable would have that and that value.

Sebastian Raschka [5:28] And so this is, in a sense, giving the model more context, more information about the training data. And it forces the model also to kind of, in quotation marks, understand training data better. So it's like, instead of just brute force, just what is the most likely next token, it has kind of like an understanding of what it is right now. I think that's also how humans work. When I, for example, as a human, read through code, I'm also trying to visualize or verbalize or write down what are the states of the variables in my for loop, for example, at this first iteration, second iteration.

Small “recursive” reasoning models (ARC, iterative refinement)

Sebastian Raschka [6:01] Back then, when I learned coding, I actually had a paper notebook and was writing down things with a pencil, basically these types of iterations. It's kind of like this approach, but for LLMs, essentially. And I do think that is something that is maybe more expensive to do, but it is also something that might push the state of the art a little bit.

Matt Turck [6:07] And what about small recursive models? What does recursive mean in this context?

Sebastian Raschka [6:23] Yeah, so that was also a big topic in 2025. That was the Hierarchical Reasoning Model. And from that, we had also another paper, Tiny Recursive Models. And so they are interesting because they were getting very good performance for their very small size on the ARC benchmark. So ARC is like a benchmark, almost like an IQ test, like a logic puzzle where there are different symbols and you have to, you see, let's say, an array of different symbols, and you have to predict, let's say, what's the missing thing here in the bottom corner.

Sebastian Raschka [7:00] And it's kind of like going a bit beyond text and beyond things usually on the internet. So, in that sense, I think the motivation behind this ARC benchmark was to have something that really tests the capabilities of that model on something new that hasn't been shown during training, like a new task, and how well the model can take some examples from that benchmark and generalize to new tricky problems. There are also different iterations of this ARC benchmark to make it harder and harder and harder.

Sebastian Raschka [7:34] The Hierarchical Reasoning Model became popular because it performed relatively well on that benchmark compared to very expensive models like Gemini, ChatGPT, and so forth. And it is a transformer architecture. And then there's the Tiny Recursive Model that is, I think, even simpler than the Hierarchical Reasoning Model. And so the idea is that you recurse. So you have a latent storage vector or something like that, where you refine the answer over multiple iterations instead of just doing a one-shot. You, let's say, write an intermediate answer, and the model looks at that: is this correct or not?

Sebastian Raschka [8:02] And takes it another round and another round and refines that answer. It is not cheap either, but the model itself is much cheaper. It made a lot of waves also because, oh, we don't need these big ChatGPT-, Gemini-type models to solve complex problems. And I think this is to some extent true, but I think also that kind of underestimates the appeal of ChatGPT, Gemini, Claude, and so forth. And the appeal there is it's like one model that can do it all.

Sebastian Raschka [8:27] It's very general purpose. You don't have to even really teach people that much how to use that model. I can ask anyone, here's ChatGPT, the interface, and people who have never used it before will be able to figure it out just by typing some prompts. And I can drag and drop an image there. I can ask it a code problem. It's doing a lot of things very well. At the same time, it's also a downside because it's this gigantic model, which is very expensive.

Sebastian Raschka [8:55] So if you have a simple task, that is very expensive to run such a big model at such a big scale. And so there is then this appeal to develop these special-purpose models. So Tiny Recursive Models, Hierarchical Reasoning Models, they are very specific to a particular task. For example, in the paper they had pathfinding, like finding the path through a maze or something like that, like a toy problem. And then the ARC benchmark, but each one was a different model.

Sebastian Raschka [9:16] It was not like one model that could do all three things. In that sense, it is, I think, really hard to compare to something like Gemini or ChatGPT. At the same time, I do think this is a very interesting and promising direction because even though, let's say, ChatGPT can do everything, it's not always the cheapest thing. If you have a business problem, you are maybe manufacturing something, maybe you can start with a generalist model, but then once you know exactly what the task is and you want to hone in on it, maybe it makes sense to replace that expensive thing by something like that that is cheaper, like a module that you can plug in.

What is a diffusion model (for text)?

Sebastian Raschka [9:46] And you can even have an LLM like ChatGPT or Gemini use those as tools. And so I think it's a great development, but I don't see it as quite a fair comparison to state-of-the-art LLMs, basically.

Matt Turck [10:01] I want to come back to something you mentioned a few minutes ago: diffusion models, especially for text. Last year, I believe Google DeepMind announced one called Gemini Diffusion. So what are those, and how different are they from transformers?

Sebastian Raschka [10:21] There's, of course, the big field of diffusion models coming from image models. Not too long ago, maybe two or three years ago, there was the big hype around Stable Diffusion, which was based on a research paper where they had a model that replaced, going back, generative adversarial networks, which were an idea for generating images. And so the diffusion models were essentially, instead of having a generator and discriminator setup, like two networks competing against each other, it was a pipeline that was denoising, starting with random noise, denoising an image, and coming up basically with realistic-looking images.

Sebastian Raschka [11:00] And you could also have a text prompt and basically guide it in terms of—it's not like a random image—you can basically guide what you want to generate. It's basically the modern generative AI image AI that we see out there. People were wondering, okay, can we do the same thing for text? So can we use this pipeline, like this denoising pipeline, to generate text instead of using transformers? Or, I mean, I'm saying instead of transformers, diffusion models can be transformers, are often also transformers, because transformers is the architecture.

Sebastian Raschka [11:34] So LLMs nowadays, it's specifically autoregressive transformers, which means these are LLMs that are generating one token at a time, where the next token always depends on the previous tokens. And so with diffusion models, you don't have that. You generate everything at once in parallel, but it might be very messy. And then you have multiple iterations. You take that whole thing and denoise it, basically refine it. What's nice about it is, well, it is fast because it's like one iteration generates something.

Sebastian Raschka [12:05] And then you have a few steps that refine that, which might be cheaper than using an LLM to generate a long response because then you have a lot of sequential steps. So let's say 16 denoising steps is fewer steps than having, I don't know, 2,000 tokens, 2,000 steps that you generate something with. The downside is, well, you have everything in parallel, and there are nowadays a lot of tasks that require sequential processing. For example, if you think about reasoning models or, for example, tool use, when you have a reasoning model and you ask it—or in general, you ask a model to answer a question—and the model maybe does a web search as part of its answer.

Sebastian Raschka [12:39] And so you have to kind of interrupt the generation. I think the diffusion models, they have these downsides, but what you mentioned is Gemini. I remember seeing the Gemini Diffusion website where they are saying something like, "Coming soon." And they compared their diffusion model to their latest, I think they call it Flash, the cheapest model. And so as an alternative to the Flash model, being even, I would say, faster at the same performance level, but they're not putting it out there as their state-of-the-art model.

Sebastian Raschka [13:10] It's more like a cheaper model, maybe for everyday use, maybe for the free tier or something like that. So it is an interesting direction to go into, these diffusion models as an alternative to the autoregressive transformers, but it is not, I would say, the replacement at the state of the art. I think one company will launch a big diffusion model this year. So there are diffusion models out there that you can use already, but I haven't seen anything at Gemini, ChatGPT, Anthropic Claude scale, I think.

Sebastian Raschka [13:23] But this year, maybe we will see something like that.

Are we seeing real architecture breakthroughs — or just polishing?

Matt Turck [13:49] Great. Super interesting. In the world of LLMs, you mentioned MoE, and that triggers a question, which is what I think a lot of people are wondering: are we seeing real architecture breakthroughs within the LLM world, or are we effectively, at this point, polishing what we already have within the LLM world? What are you seeing that's moving the needle in terms of architecture improvement or optimization?

World models: what they are and why people care

Sebastian Raschka [14:13] Improvement is not so much coming from the architecture anymore. It is basically the post-training. But coming back to the architecture, I think it's still an interesting question because there are so many different architectures, and almost no one uses the same one. They're all very similar, but they are not identical. I think a lot of it is coincidental, where there are some tweaks. And if you look at the loss, in some cases, on some training data and some training pipelines, maybe moving the normalization, the RMSNorm, before or after makes a small difference.

Sebastian Raschka [14:48] I mean, there are theoretical justifications. But also, for example, OLMo 3, which is very transparent, they moved the RMSNorm placement. So then Gemini had a post- and pre-norm. They had both on both ends. And so there is some justification where, okay, ablation studies show this stabilizes the training. But while assuming stable training, it's not going to, I think, make your model magically perform better. I mean, this is just like people tune their cars a little bit by putting in different air filters and something like that.

Sebastian Raschka [15:18] So I think it's on that level where you make small tweaks, but it's not really changing the engine itself. The one thing, though, is what we've seen: a lot of large architectures now using MoE. I think that's a new 2025 thing. Of course, MoE was not invented in 2025. That was, I think, going even back to the Google Pathways paper in, I don't know, 2022, '23, something like that. And then Mixtral had a big MoE, I think it was 2024.

Sebastian Raschka [15:40] Then I think it was pretty quiet. I mean, around MoEs, there was only, I think, the ChatGPT model, which was rumored to be an MoE. But now this year, really almost everyone has an MoE out there, like every open-weight developer. I would say DeepSeek kind of restarted that trend in 2024, in December, with DeepSeek V3. They had an MoE model before, but I think this is the one that everyone looked at because that made such a big splash that people were like, oh, what they are doing is maybe sufficient.

Sebastian Raschka [16:14] It's the right thing. Let's not try something crazy. Let's iterate on that. So there were a lot of companies adopting straight-up the DeepSeek architecture. So there was, I think, Kimi had the DeepSeek architecture, scaled it up, I think, from 670 billion to 1 trillion parameters. And then even the European Mistral AI company used the DeepSeek V3 architecture for their new Mistral 3 model. It is something that is working well, but then DeepSeek itself, they iterated on that too.

Sebastian Raschka [16:34] Where they changed the attention mechanism. They had a multi-head latent attention, which is already a nice tweak to—they added sparse attention, where sparse attention, again, it's not new, but they had their own flavor of it to make it cheaper. The idea, I think, is to get better modeling performance through the training pipeline while tweaking the architecture so that benefits can be, of course, absorbed by the architecture, but then also, at the same time, to bring down the cost of running the architecture.

Sebastian Raschka [17:06] GPT-4.5 was rumored to be a bigger model, a bigger version of GPT-4, but it was not very popular because it was too big, too expensive. And so they kind of abandoned it and went a different direction with GPT-5. And so I do think, well, I wouldn't expect bigger architectures. I would expect more efficient architectures, tweaks, getting the same modeling performance for less compute, because then you can have more tokens for the same cost, and the tokens, they give you better performance, like inference scaling and so forth.

“Pre-training isn’t dead… it’s just boring”

Matt Turck [17:32] But you see room for progress there? Like, you're not in the pre-training-is-dead camp?

Sebastian Raschka [17:53] I would say pre-training is not dead, but pre-training is boring. So it's not where the low-hanging fruit is anymore. I think the low-hanging fruit used to be in pre-training, but now you need really good pre-training still. But it is, I think, harder. I mean, I wouldn't say harder, but you can get better bang for the buck elsewhere, almost, I would say. Pre-training, I don't think it's dead. It's just not, let's say, the most popular thing to spend money on right now.

2025’s headline shift: RLVR + GRPO (post-training for reasoning)

Sebastian Raschka [18:03] I think it would make more sense to put a lot of that budget into post-training right now.

Matt Turck [18:32] Okay, let's go into post-training. So in your blog post, you mentioned that 2025 was the year of RLVR and GRPO. So you had a nice timeline where you said 2022 was RLHF, which gave us ChatGPT plus PPO. '23 was LoRA SFT. 2024 was the year of mid-training, and 2025, the year of RLVR and GRPO. So I would love it if you could walk us through those techniques. So fair to say, both of those belong to the world of post-training.

Matt Turck [18:46] Let's pick RLVR, and let's start with the definition. What does RLVR mean versus regular RL?

Sebastian Raschka [19:13] I would say RLHF is the biggest leap in LLMs we have seen in a long time, because that was taking GPT from GPT to ChatGPT, like the RLHF, the reinforcement learning with human feedback. And in that sense, it's almost like RLVR, which is reinforcement learning with verifiable rewards, took that other leap, basically from just a simple chat model to a reasoning model. Both RLHF and RLVR have the RL in it, so both are based on reinforcement learning. But, I mean, this reinforcement learning is a bit different from the reinforcement learning that plays Go.

Sebastian Raschka [19:41] It's almost like it's a special thing and a simpler thing in the context of LLMs. But the idea is that instead of doing next-token prediction, just predicting what's the next token, it's more like looking at the full answer. And then, based on that answer, you give a reward. Like in RLHF, you have multiple answers and then you say, which do you prefer? Or in the case of RLVR, you look at the full answer and then let's say it's a math problem.

Sebastian Raschka [20:07] You say the math problem is correct. The final answer of the math problem is correct or incorrect. That's like the main difference between next-token prediction and pre-training, and then the RL here. So RLVR was kind of popularized by DeepSeek R1, which was based on DeepSeek V3. And that came out—R1 came out in January 2025. And with that, they also introduced the GRPO algorithm you mentioned, but they go well together because they make the whole thing more efficient, but it doesn't have to be.

Sebastian Raschka [20:38] So you could technically do RLVR with a PPO algorithm that was used back in RLHF. Now, why I think it's such a powerful combination is, well, it just makes things more efficient. With RLHF, you had to have people ranking answers because the goal was essentially to train a model that prefers one style over the other. So, for example, for safety, like reducing swear words, if there are two, use the one with fewer swear words, or if you have an explanation, maybe use the explanation that is simpler to read, and these types of things.

Why RLHF is expensive (reward model + value model)

Sebastian Raschka [21:17] But you always have to have someone who compares these answers and says, okay, this answer is better than the other answer. What you do then, though, is during RLHF, you train a reward model, another LLM that provides this information for you. So at that point, you can replace humans looking at these answers. So you have this other model that does it automatically as part of your loop. It's more expensive. Now you have two models, essentially. And then there's also a value model.

Why GRPO makes RLVR cheaper and more scalable

Sebastian Raschka [21:45] So the value model is internally kind of like a reward model, but it also gets updated to make some predictions as part of the reinforcement learning signal. And so you have basically three models in memory. And if you have ChatGPT-style training, like large models, or even like DeepSeek-V3, 600 billion parameters, you have three times 600 billion parameters, and you have to keep them all in memory. It's very expensive. And so in RLVR with GRPO, you replace two of these models.

Sebastian Raschka [22:09] So you have three models for RLHF with PPO. You replace that reward model with verifiable rewards. So instead of having someone say, "Oh, I prefer this answer over the other," or using an LLM for that, you now have tasks that can be automatically verified. So, for example, with math, you can have a math parser. It could be something like Wolfram Alpha. You have the correct solution and the LLM solution, and you just parse out that part that you can compare algorithmically.

Sebastian Raschka [22:38] And then based on the correctness, you can give a reward for the reinforcement learning. So you already eliminate one big LLM that you have to train and have in the loop. And the other one you also eliminate. So there's the value model that assigns a value to each of the responses during the training.

Matt Turck [22:38] Yeah.

Sebastian Raschka [23:05] Instead, you just compare them relative to each other. That's where the R in GRPO comes from, like Group Relative Policy Optimization. And this makes it much more feasible to train. It's just cheaper. And they show that it is actually really powerful. So you can take a base model, even skipping supervised fine-tuning in RLHF, and just do this RLVR, and you get a really good reasoning model out of it. The DeepSeek-R1 model—you can still do supervised fine-tuning in RLHF, and it's recommended to do it.

Sebastian Raschka [23:33] Reasoning behavior comes from that RLVR. There are, of course, papers showing that the base model already has reasoning capabilities. And I think this is actually partly true. But it is hard to say for sure because you don't really know what's in the pre-training data anymore. So there's also a lot of reasoning data. So reasoning data is essentially just data which has this chain-of-thought format, which means that the model writes intermediate steps, like it explains its own answer.

Sebastian Raschka [24:07] A lot of the pre-training data already has that style of data in it. And then it's hard to say: does the reasoning behavior come from the pre-training corpus, or is it from the RLVR? And in my experience, I think a little bit of both. So, for example, I took the Qwen3 model as part of my book, the Reasoning from Scratch book, and I trained it just for 50 steps with RLVR. And it goes from 15% accuracy on MATH-500 to 50% on MATH-500.

Sebastian Raschka [24:40] So it takes this threefold leap in terms of accuracy by only doing 50 reinforcement learning steps. And I think it's not really learning that much in these 50 steps in terms of how to do math better. I mean, it does, but it's not learning new knowledge about math. The knowledge is already there in the pre-training, and this just unlocks it. It's just like a step that maybe shows the model how to use its own knowledge, basically. So that's how I think about it.

Sebastian Raschka [24:43] Yeah.

Process Reward Models (PRMs): why grading the steps is hard

Matt Turck [25:00] Fascinating. Just to unpack some of this, you can understand the reasoning steps that led to the explanation. You mentioned in some of your writing the label process reward models, PRMs, and the fact that this is not successful yet. Can you unpack that part?

Sebastian Raschka [25:27] So there's an outcome reward and a process reward. And the outcome reward is mainly like, is the final answer correct or not? But then there's the whole explanation of the reasoning model, whether it leads to the correct answer. And so there's also research like, hey, why should we throw out everything the model generates and only look at the final answer? Can we get something useful out of this intermediate explanation? And the intermediate explanation is useful for several reasons. I mean, one is it has been shown that this helps the model to generate the correct answer, whether the explanation is correct or not, but it's a different aspect.

Sebastian Raschka [25:57] But just the fact that it generates these intermediate steps is correlated with a more accurate answer. Then the hypothesis is, if we can improve that explanation, maybe it gives an even better answer, like maybe it even drives the accuracy higher. If you want to learn something, it's not enough to just see the final answer. You want to see the steps that lead to the final answer. Process reward models are also focused on training the model to reward the models based on that explanation.

Sebastian Raschka [26:24] And so my statement that it is not so, let's say, promising or useful was mainly based on the R1 paper, where they had a final paragraph at the bottom. I mean, this is already a year old, but they had a paragraph at the bottom that I think the headline was something like "Unsuccessful Attempts." They tried it and they found it wasn't worthwhile because of reward hacking. Because usually you need another model to grade the responses, and they can be susceptible to reward hacking, and it's hard to train that model, and it's not reliable.

Sebastian Raschka [26:54] And then that whole thing, it's not really worth it according to their experiments. There are a lot of people who try to make it work. And I think it is promising, and we will see it working at some point, I think. So it's just like right now, it's still tricky to make it work, but I'm quite sure we'll see it as part of the standard repertoire at some point. And there was, for example, at the end of last year, the DeepSeekMath-V2 paper.

Sebastian Raschka [27:19] Where they had actually a nice study. They had something like that where they had a second model that was checking the answers and explanations of that first model. It's almost like turtles all the way down. They had yet another model. So they had three models. They had one model generating the answer, one model to grade the answer and the intermediate steps, and they had one model to grade the grader, basically to say, oh, is the grader actually doing a good job?

Sebastian Raschka [27:47] So there were like three models in a row. It sounds a bit excessive, but based on the performance of that whole setup, it performed really well. So they were cranking up also the self-refinement steps and iterations. And they got gold-level performance on some of the math benchmarks. One could say, okay, maybe, well, cheating, the data was public or whatever. I don't know, that's a different question. But the fact that this performed better than the model without it tells me whether, let's say, it's really gold-level performance is a different question.

Can RLVR expand beyond math & coding?

Sebastian Raschka [28:20] But it is doing better than just the plain model, so it is actually adding usefulness to the whole process. And I think we will see more of it. It's just expensive because now you have to have more models, more training, more stuff. But that's what I meant earlier with that's where you make the bigger gains rather than scaling the model size. I think that's one of those things where you will see more progress coming from.

Matt Turck [28:32] And speaking of math, I think that's one of the key questions going forward for RL, whether you can expand this beyond math and coding to other domains. What's your take on that?

Sebastian Raschka [28:56] Yeah, I think what's so attractive about RLVR is that you don't have to have, let's say, humans checking the solutions. You have a verifier that deterministically checks for math: is the answer correct? Like, given two fractions, are the fractions the same? Or two numbers with decimal points, are they the same if I round them up? And so it's very easy to check programmatically, algorithmically. And the same for code. So you have code problems, and in that case, you can compile the code.

Sebastian Raschka [29:24] If it compiles or you have unit tests, it checks it works. It's very nice to check. There's no subjective aspect. It's very objective. You can say, okay, it compiles, it doesn't compile. It's very clear-cut. The question you had is: what happens now, or in general, to other fields? Is it specific to math or code? And I think we will see expansions of that to other fields. I'm personally not an expert in other fields, so I don't know what that would look like for medicine.

Sebastian Raschka [29:53] I have a computational biology background. I know a little bit about the drug development pipeline and so forth, but I think it's not quite as clear what the reward looks like. But you can also be more creative. It doesn't have to be strictly verifiable through an algorithm. It can be verifiable maybe through an LLM. It could be something like that. So, for example, I can see, I don't know, for research, maybe training a model to give correct citations, and you can maybe check the citations.

Sebastian Raschka [30:21] You could have another model that goes through the URL and says, oh, this is indeed the correct paper, giving the correct title of the paper or something like that. It's the correct URL and stuff like that. So I think there are lots of these things where we can expand RLVR to and train on those things. So I think we will see a lot of that.

Why RL feels “finicky” at scale

Matt Turck [30:46] The thought that crosses my mind as you describe this is that I've heard people say that RL is very difficult to scale, is very finicky. And hearing what you describe about different techniques put together, I'm starting to get a sense for why. Is that why it's complicated? Basically, a bunch of different things and models talking to one another, that it is hard to scale, or is there another reason?

Sebastian Raschka [31:07] Well, I just implemented, before we were recording this, GRPO RLVR from scratch in a Jupyter Notebook. I wrote up a chapter, 39 pages. So it is not super complicated, I would say. You can fit it into a Jupyter Notebook. It works and it trains fine. What I'm trying to say is, if you can figure out pre-training, the scale at pre-training, you can figure out this. Because if you also look at the numbers of how much it costs, just GPU hours.

Sebastian Raschka [31:35] DeepSeek V3, they had like a $5 million price tag on that, given the, I think, $2 per GPU hour they assumed. Whether that's a correct assumption or not is a different question. But if you compare it relative to the cost of R1, I think R1 was about $300,000 when they trained it. They had a number in the Nature version of the paper. So it's basically more than 10 times cheaper than pre-training. And it's the same infrastructure where you have to make sure that GPUs don't crash. If they crash, that you can resume, and so forth.

Sebastian Raschka [32:07] And also during pre-training, you might have bad losses where you want to roll back to the old checkpoint. And the same things apply. You have multiple models, but yeah, you are right. There is a bit more, I would say, trial and error in RLVR, where, just due to the nature of the updates and so forth, that's what I observed when I was training my models. You often get—not that often, but every so many hundreds or thousands of steps—the model gets bad.

Sebastian Raschka [32:28] So the model works totally fine, and you train long and long, and suddenly the model is really bad. And so you just go back to the previous checkpoint. But it's not new in terms of a new thing that's happening all the time in pre-training as well. But I think vanilla GRPO, the original algorithm, is pretty flaky, where you have to babysit it. Over the course of the year, many people had these tips and tricks where some people were saying, remove the KL divergence term.

The practical “tips & tricks” that make GRPO more stable

Sebastian Raschka [33:05] If you just drop it for math, it performs better. Remove the standard deviation normalization term. Or if all the rewards look the same, you can skip them to make it faster. There's a lot of tips and tricks, these tricks of the trade that make it more stable. And I think if you apply all of them together, it is actually a pretty stable, okay-ish algorithm. Just like last week, NVIDIA also had a paper on GDPO, I think.

Sebastian Raschka [33:35] Yeah, GDPO. So they were focused on algorithmic improvements with respect to multiple rewards. So if you have more than one reward, it could be something as simple as you have the accuracy reward, but you usually also have a format reward because you want the model to put the final answer in—you don't have to, but you can put that into these think tags. So there's a think token. It's more, honestly, for stylistic purposes. I think the advantage is some people develop models that are hybrids, which are capable of a normal mode and a thinking reasoning mode.

Sebastian Raschka [34:06] So thinking stands for reasoning. The appeal here is you don't always want to use reasoning modes because it's expensive. It uses a lot of tokens. And sometimes you have a simple answer and you don't want to spend 2,000 tokens on the simple answer. And so you can, with these think tokens, steer it a bit. For example, in Qwen 3, you can add empty think tokens. So you have an opening token and a closing token. If you add that think, whatever in between is empty, and then the model will not generate any reasoning chain of thought.

Sebastian Raschka [34:40] And long story short, during training, you can teach the model to adhere to these different formats. So then you suddenly have a second reward. So one reward is correctness: is the answer correct? The second reward is, does the model output something that fits my formatting here? And then you have two rewards, and how you combine them. Usually, originally, you just add them up together, but then there were some downsides in the GRPO instability. And so GDPO had some algorithmic improvements to improve the stability.

Sebastian Raschka [35:07] And so there are lots of these little tricks over the years to make RLVR more stable, but it is a newer paradigm. So it just takes a few iterations to find the canonical one. It's similar to optimizers with Adam. So Adam is—I mean, right now there's AdamW, there was SGD and all the other RMSProp and what they were called. And they kind of all converged to AdamW by adding more and more tricks. And I think that's the same right now with RLVR, with GRPO, we're adding more tricks.

Sebastian Raschka [35:17] To get to something that is pretty stable across a lot of different scenarios.

The meta-lesson of 2025: progress = lots of small improvements

Matt Turck [35:40] It's fascinating to hear you talk about tips and tricks and different techniques. That triggers the thought that you had a nice way of putting it in your blog post, and taking a step back for a second from the weeds. You talked about a meta lesson for all the things in 2025 and where progress actually comes from. Do you want to get into that? I think that'd be interesting.

Sebastian Raschka [35:58] Yeah, and so the meta lesson would be essentially that, well, I think we are talking right now for half an hour about different things. So I think the theme would be, well, there's no one thing that fixes it all. It's a lot of little tips and tricks all over the place. And if you add them up, that will give you the progress. But I think there's no magic lever, no magic, I guess, bullet that gives you everything.

Sebastian Raschka [36:26] It is kind of tweaking things here and there and making things more robust. I think the tweak was the Transformer architecture back then. And now it's essentially, let's make it even better, I guess, refining it, and a little bit of post-training here, a little bit maybe improving the quality in pre-training, maybe some architecture tweaks, algorithmic tweaks. It's all a little bit of everything, basically. It's also that I would say, you don't have to know all these things because, in practice, it's like a big team at a big company.

Sebastian Raschka [36:48] And everyone has a specialty, like everyone is either on the post-training team or the pre-training team. It's not that one person has to know everything and all the tricks, because that would be really impossible. And so I think it's also just due to the nature of work, because it's so much work. It's a lot of work to train these big models that you kind of separate these roles, and then everyone can work on everything at the same time, which is also nice.

Sebastian Raschka [37:02] And then you bring back together all these improvements into the model. Yeah.

Matt Turck [37:08] And you're confident in the industry's ability to keep coming up with tricks and tips going forward?

Sebastian Raschka [37:34] Yeah, that is a good question. I mean, if I look at DeepSeek, for example, because I'm always picking DeepSeek in this podcast because I think they have a really nice trajectory of models. I wish I could also talk more about Gemini and ChatGPT, but they don't really release the techniques, so it's hard to talk about it. The DeepSeek V3.2 model with the sparse attention mechanism, and then also this DeepSeekMath-V2 with self-refinement and everything. So they do have, right now, still a track record of improving things, and they are rumored to release a new model in February, the DeepSeek V4.

Sebastian Raschka [38:09] But I think, well, so far, yeah, I think we are still on that trajectory where we haven't run out of ideas. So I think the only thing we are running out of is really benchmarks. So the improvement on benchmarks, it's kind of harder to measure. And I think maybe it's not the one-shot problem anymore, where it's not really answering knowledge questions. It's not really solving math problems in one iteration of the benchmark. It is maybe more like the—

Matt Turck [38:17] Yeah.

Sebastian Raschka [38:40] Agentic cycle, where you have more of an objective that is not, let's say, answer the question, but more like design something, blah, blah, blah. And then it goes off, and how long it can, or how long it needs, or how long it can run until the problem is solved. And I think it's maybe more towards that, how we measure progress, rather than whether we get 90 or 95% or 97% on a benchmark.

“Benchmaxxing”: why benchmarks are getting less trustworthy

Matt Turck [38:48] Yeah, yeah. You had this nice expression, benchmaxxing. You mentioned benchmaxxing in some of your posts. What do you mean by that?

Sebastian Raschka [39:13] Yeah, so benchmaxxing is—so I'm often reading things on X because that's where a lot of the AI community is. I think benchmaxxing is one of the ones that came up in 2025, like a newer-generation term. And so, loosely, what that means is essentially that, well, it's almost like exploiting the benchmarks: do well on the benchmarks, but it doesn't really translate to real-world performance. A popular example was the Llama 4 model. I mean, based on rumors I heard, they had a separate model just for the benchmarks, the leaderboards.

Sebastian Raschka [39:36] But let's say even that aside, if someone, let's say, trains a model on leaderboard performance, it doesn't mean the model necessarily performs better in real life because leaderboards are also susceptible to style. And so, with leaderboards, the tricky part is because humans compare which model they prefer. And if I have, let's say, a very complicated math problem, I ask an LLM, and let's say I don't even know the answer, and then I—well, it should help me with my tax report or something like that.

Sebastian Raschka [40:12] And there's one LLM that gives me a really nice explanation. Maybe the result is wrong, but the explanation is really nice, easy to follow. I probably like that one. And so I would probably give it a thumbs-up because, oh, it's understandable, it's reasonable, because I don't know what's correct because I'm not an expert. And I think that's one problem with leaderboards. It rewards the style more than the correctness because there is no correctness check. It's like, yeah, you, as the expert, have to know whether it's correct or not.

Matt Turck [40:17] Yeah.

Sebastian Raschka [40:40] And so also, LLM developers, when you're training the LLM, it kind of gets biased to follow a certain style, and the style of people who use those leaderboards. And in that sense, you end up with models that have, let's say, been benchmarked. They have been getting really good benchmark scores, but they might not do better than previous models. And then it's kind of like a tricky thing. It's hard to measure progress this way.

Matt Turck [40:52] Do you think people do that just out of largely economic incentives? Like, the companies need to raise more money and people need to have successful careers, and therefore they want to look good. Is that the driver?

Sebastian Raschka [41:16] I mean, I don't want to accuse anybody. I don't know for sure. I mean, I only know what is known on the internet. I read on, let's say, Reddit a few times that Llama 4 was a separate model. So there might have been some company-level decisions that have led to that. I honestly don't know. And maybe incentives, getting good headlines and that stuff. But, well, I think the open-weight community is a pretty smart community, so it's like, I think it's not worth risking something like that.

Sebastian Raschka [41:42] And I think most people don't risk it. It's just implicit. It just happens. It's like if you iterate too many times, it's a classic deep learning problem or machine learning problem. But the nice, the beautiful thing here was actually it's not a big concern because it happened to all the models. So all of the models performed like 5, 10% worse on this new data. It was pretty consistent. So if you were to rank those models, the ranking would still be the same.

Sebastian Raschka [42:07] So in LLM terms, let's say ChatGPT and Gemini, let's say they achieve on the benchmarks, and the models are 10% worse. But if the ranking is still the same, let's say Gemini is still better than GPT, then it's not a problem if both of them do that. And I think we have right now that in LLMs where I wouldn't say they are cheating, they are just using the data a lot. And from using the data a lot, well, the data leaks in a sense.

Sebastian Raschka [42:32] So you're kind of, like, biased in a sense, but then if they're all biased, then it's again fine because the ranking is still the same. But I think, yeah, the problem still remains. The benchmarks are saturated, and it's hard to demonstrate or detect or have any type of notion of progress. It's really right now, honestly, personally, I stopped looking at the benchmark numbers. I just use the model and see for a few days, and I see if it's better or not.

Sebastian Raschka [42:58] Like, I can't say, okay, this is better by so-and-so many percent. It's more like, oh, I use it and I feel like I can't even put it into words. And I think that's the challenge we have right now. How do you put that into words to communicate the progress? And I think that will be, in the upcoming years, the more difficult problem to solve: how to actually evaluate what you're using. I mean, the power of LLMs is that they are so freeform, but that's also the downside for evaluations because evaluations, if you want to be numeric and precise, well, freeform is not so easy to deal with.

The other big lever: inference-time scaling

Matt Turck [43:36] Super interesting. So, going back to tips and tricks to make sure that we cover the state of LLMs in, I guess, early 2026: we talked about post-training. How much of the recent progress in the last year or so do you think comes from non-architectural and non-post-training stuff? And I'm thinking in particular inference scaling and then tool use.

Sebastian Raschka [44:01] I do think a lot. I mentioned previously post-training, but honestly, I think one of the biggest drivers this year has also been inference scaling. So inference scaling essentially means you don't change the weights of the model; you just expend more compute while using the model, when the consumer or the user uses the model. And a beautiful example or chart was back in October 2024, when OpenAI o1 came out. They had this chart where they had two subgraphs.

Sebastian Raschka [44:30] One was for scaling the training, and one was for scaling the inference. And you could see for both, both were going up with a similar increase. And so you can basically invest either more money during training, which is a one-time cost, and then you have a fixed-size model, and you never have to pay money again later on. But then that breaks a bit, in a sense, because reasoning models generate more tokens, so they are also expensive to use during inference. So inference scaling includes that.

Sebastian Raschka [44:56] It includes models that generate more tokens, because if you generate twice as many tokens, it's twice as expensive because now you have twice as many steps. That is one form of inference scaling, but it can lead to more accurate answers. The other one is parallel sampling. You just have the model, you ask the model multiple times, more like a majority vote. And most people, if you see the benchmarks, they do that. It's called best-of-N, something like that, like best-of-five or something, or best-of-five, I think, running it five times and then selecting the answer.

Sebastian Raschka [45:31] So you can do majority vote, but it's five times more expensive now because you have to run the model five times. There are also methods where people have a judge model that judges these results. Like, if you can't do majority voting, you can have a score and then score the highest answer. It's a bit brittle because that model can also make mistakes, but there are all these types of tricks, or self-refinement, where you have multiple iterations. Self-refinement is also basically these iterations where you have one LLM write the answer, and then you say, okay, take a look at this answer, and then, oh, I made a mistake here, and it self-refines.

Sebastian Raschka [45:57] I mean, reasoning models do that internally as a chain of thought also sometimes, but you can also have an explicit version of that. A really cool paper that came out in January was RLMs. And so what they do is they take that prompt. So instead of processing it all at once, they chunk it up into several smaller prompts, or the LLM decides, it learns how to, or sees how it should chunk it up in code, and then runs a prompt on each of those again.

Sebastian Raschka [46:32] So basically making one prompt into smaller prompts and then having multiple requests. And this, I would say, is also a form of inference scaling because it depends, but I do think it can be more expensive because now you have more LLM calls and each one, if you want to go deep, you can end up spending more tokens. Not more tokens in one request, but in the sum of all the requests. But there are all these things, I think, that are underappreciated in a sense because they're not—I mean, in this case, it was a popular paper, but often inference scaling is not talked about that much.

Sebastian Raschka [47:04] But I think it is a big driver of making LLMs perform well. And I think why I think that is, if you use DeepSeek locally or use the platform—let's say you use a local LLM and use ChatGPT—I think ChatGPT, of course, is a really good model, but I do think the leading open-weight models are not that far behind. But if you use them locally, they don't feel as good. And I think that's because ChatGPT has a really good interface, like the platform they have.

Sebastian Raschka [47:30] It's not just running the LLM; it's maybe cleaning up your prompt. Again, this is, I would say, a hypothesis, or I'm guessing here, but instead of having the LLM learn—of course it can and it does learn how to deal with misspelled words—you can also just clean that up, the input, in certain cases where that might improve the accuracy. And I think all these little engineering tricks, not just inference scaling, but just cleaning up the prompt, how to manage the context, the history, and everything, I think that all contributes to a lot of progress that is felt by the user.

Tool use: reducing hallucinations by calling external tools

Sebastian Raschka [48:08] Yeah, and then another example would be tool calling too. I think that was also a big one in 2024. I don't remember when ChatGPT introduced tool calling, but it might have been early 2025 or 2024. But GPT-OSS has—so they had that open-source model in, let's say, summer 2025. And GPT-OSS has tool-calling support. So tool calling means that the LLM can call a web search or it can call code interpreters and so forth. And that is very, very powerful because I think this is one of the ways you can mitigate—not totally mitigate, but let's say reduce—hallucinations, because then the LLM suddenly doesn't have to remember everything anymore.

Sebastian Raschka [48:34] You can outsource a lot of things that are hard to tools. So, like we humans do, right? We humans use calculators, we use web search, we don't try to memorize everything. And so I think, in that sense, that is really a big unlock. The only problem is, well, you have to trust the LLM to run on your computer, which is why right now I think it's more confined to these proprietary LLMs like Gemini and ChatGPT because, well, it's not your computer it runs on.

Sebastian Raschka [49:12] If it goes somewhere, executes some code, and messes it up, well, not your problem. But I think we will see more of that in the upcoming years when the open-source tooling kind of, I would say, gets more robust and people have more trust in running that on their own computer, maybe in a Docker container still, but yeah, something like that. I mean, right now a lot of people already run code agents on their computer. They're usually kind of more restricted, but I think people are more and more trusting these to do things they wouldn't have trusted them to do, like, a year ago, and give them permissions.

Sebastian Raschka [49:46] But as LLMs get better, people develop more trust—or, I mean, run it in a secure virtual environment. And I think that will be a lot of small leaps we will see, even though the LLM doesn't get bigger or anything like that. And also, you can actually go to the GPT-OSS release blog, and they did have benchmarks to show how the performance on the benchmarks is with the same model with tool calling enabled and disabled. Two times the performance. But you can see there's definitely a jump in capabilities just by allowing the model to use tools, basically.

The “private data edge” + in-house model training

Matt Turck [50:11] And that's part of where you see the world go, right? This combination of open-source models and private data. I think you call that the edge in your blog post.

Sebastian Raschka [50:36] What distinguishes different LLMs right now? They're all kind of similarly good, I would say. Open-weight LLMs—a lot of LLMs are similarly good. I mean, personally, I don't use all of them all the time, so I usually use one LLM at a time. But if you use or compare ChatGPT, Gemini, Claude, Grok, I think they're all pretty much on the same level. And I think that's because they're trying to do everything. They're generalist models for a general person to do a lot of things.

Sebastian Raschka [50:59] I mean, Claude is a bit more specialized to code now, but the other ones are more like general models. And I wouldn't say one is significantly better than the other. They have small differences and so forth. But if you want to really distinguish them and make them better in certain industries, I do think the private data is what helps, like all the treasure troves of data that a finance company has collected over the years, over 100 years or 50 years, or medical data, like medical records from patients.

Sebastian Raschka [51:36] I think ChatGPT had a contract now to process them, to make it secure and private. But I honestly think these companies don't want to just give away that data. First, they can't. I mean, it also makes sense. As a patient or a customer, I would feel really bad if someone gave my health data to some other company that I didn't agree with, or agree with this sharing. But then also, the companies don't want to just give away all that data because once they do, maybe they become really kind of obsolete in that sense, where all the treasure is basically all there.

Sebastian Raschka [52:09] What makes them different from other companies is now taken, basically. And so, over the last month, people reached out to me also. I know for a fact that big companies are training LLMs in-house. Really big companies who have the financial means to train ChatGPT-like models are hiring people who train LLMs. And I think that is also what we will see: that instead of going to these big LLM providers and giving them the data, people will try to make their own LLMs for their own company and private data.

Matt Turck [52:36] It's fascinating, right? It's sort of back to the future because initially people thought that they were going to train the models, and then they kind of gave up. But what you're saying is that you're seeing people going back, maybe with a better state of open-source LLMs that they can build on as a building block. That's what you're saying, right?

Sebastian Raschka [52:55] Yes and no. I think you're right. No, you bring up a good point. Open-weight and open-source models are very, or were very popular a few years ago and are still very popular. And I love working with them. But I think there's still a gap between a ChatGPT model and an open-source, open-weight model. Maybe now with DeepSeek version 3, not so much, but that's almost a different community, like the tinkerer community, like me, like small systems.

Sebastian Raschka [53:29] Well, DeepSeek version 3 would be way too expensive to run for me every day. I would have to spend thousands of dollars on just hosting costs every day, every week. And so I toy around with smaller special-purpose models. But what I meant is, first, the open-source community in that sense will maybe have a comeback at these companies. But I even mean a step further: that they actually develop models from scratch, like really big models. And what's different from, I would say, the regular open source here is that it is really large-scale.

Sebastian Raschka [54:00] It's like a ChatGPT-scale data center, large LLM. It's not something you run on your computer, basically. It's really like a big data-center-style LLM. So I know that there are—I can't say any names—but I know people are interested in that. They are exploring that. Whether it will work out or not, I don't know. But I think right now, if you are in college, you are learning about LLMs. That's the thing.

Sebastian Raschka [54:24] Big thing. So you start with open source, you start training small LLMs, and you work your way up. And then at some point, you probably want to get hired either by Gemini, ChatGPT, or so forth and do the big model development there. But not everyone—I mean, you can't have 100,000 people doing that. So there will be people distributed across different companies who will do something similar. And also, on the other hand, people who are, I think, at OpenAI and Gemini at some point—I mean, big finance companies have a lot of deep pockets—they will make it attractive to do something similar at their company too.

Sebastian Raschka [55:04] So I think we will see—right now it's very concentrated at these companies—but we will see the knowledge being a bit more spread out, where other companies will develop models too. We will probably not hear about it in the news or anything. I mean, the news, of course, but there won't be papers, there won't be big announcements because it won't be this general-audience model. So it's more something they will do internally, right?

Continual learning: why it’s hard (and why it’s not 2026)

Matt Turck [55:33] And especially if you're a big hedge fund or a defense company, I assume that's the kind of companies we're talking about. Yeah, those have a very secretive DNA. Okay, fascinating. What else do you see happening this coming year? I was interested to read your thoughts on continual learning, which was sort of the talk of the town at NeurIPS, but you viewed it, or you view it, as a 2027 thing, not necessarily a 2026 thing?

Sebastian Raschka [56:00] Yeah, continual learning is an interesting one. I think it was discussed there very heavily, but also in general, if you went to social media, AI-related topics, continual learning was there all the time, everywhere. And to be honest with you, I don't know why exactly it was such a hot topic in 2025, because I don't think there was a big breakthrough. Maybe it's the hope for a breakthrough, or the, like, hey, nothing has changed, so maybe we have to focus more on that to force some change.

Sebastian Raschka [56:27] But yeah, continual learning, I think, is an interesting topic because it sounds attractive if you have an LLM that self-improves, or an agent that does something, fails, and learns. I don't think anything like that is feasible this year. Right now, you could technically do continual learning if you wanted to with the data you have, like when you think even of RLVR. So you could technically—it's just updating the model—but the problem is still catastrophic forgetting.

Sebastian Raschka [57:00] Also, you don't want to train the model on garbage data. So people kind of do—I would call it continual learning—but in a more controlled setting where, instead of just updating the model, letting the model update itself, they collect failure-case data and then construct the dataset and do it in a more controlled manner. But you can see, based on the model releases, it happens more frequently than it used to. Because back then it was GPT-1, GPT-2, GPT-3.

Sebastian Raschka [57:29] Too, and all the models, they are iterating now, or even the same with DeepSeek. So it's like this same architecture, you iterate multiple times, but still in a more controlled way. And I think it makes sense because it's such an expensive thing to do. I don't even know how you would do continual learning if you have this model that is hosted in a data center. You can't just update it. It's a very expensive model. You can't just update it on good luck.

Sebastian Raschka [57:52] So it's a risky thing to do, and you can't do it even as a single person. You have to be really careful, monitoring a lot of things. I don't know how that would work because right now we are still in this era where everyone uses the same model. People don't have custom models. When I go to ChatGPT, I have the same model as you do. And yeah, the prompt is a bit different, like the memory and everything, but it's all in the prompt.

Sebastian Raschka [58:26] It's the same model weights. And as long as we have that, I don't think we will see anything like continual learning. There are companies, I guess, like Tinker API. It is something where it is democratizing the training a bit, where through an API you can now train your model more cheaply, or instead of having the hardware, it's on their data centers. People have their own copy, but it's very, very far from continual learning. It's just making available what other companies have in terms of training on a large cloud instance without you having to set it up. But I don't see anything in 2026 that really makes continual learning the big breakthrough in efficiency, or I don't know.

Sebastian Raschka [59:06] And so 2027 is even ambitious. So maybe we'll see something there in 2027. Given that it is such a big topic, an important topic, there's a lot of smart people thinking about this and working on it. So we'll probably see something, or at least some ideas that are fresh or prototype things that work or are interesting, but it's really hard to say anything concrete without having seen anything. So it's a prediction that I just put out there.

Sebastian Raschka [59:14] Maybe we'll see more continual learning stuff in 2027, but with a grain of salt.

How Sebastian works: reading, coding, learning “from scratch”

Matt Turck [59:52] Yeah, maybe it becomes a self-fulfilling prophecy, right? If you have enough smart people that decide it's the thing, then maybe it does happen. All right, so maybe to close the conversation, let's talk about your work and how you do your work. So you published a book this past year, in 2025, on how to build an LLM from scratch. I believe that you are writing the sequel currently, How to Build Reasoning Models from Scratch. Is that the title?

Sebastian Raschka [59:52] Yeah.

Matt Turck [1:00:12] And you produce an incredible amount of work. So people should find you on your website, on your Substack newsletter. How do you absorb all of this knowledge? To which extent are LLMs part of your workflow? I'm just curious how you work these days.

Sebastian Raschka [1:00:34] Good question. So I must say, I don't have a magic approach or anything. I think the thing I have maybe is I get very excited about things. And then when I'm excited about something, it goes very easily and very fast. I don't know. If you notice, maybe I write only about certain topics. I don't cover image models at the moment, for example, because I am just very excited about LLMs. And then, I don't know, I just can't help it.

Sebastian Raschka [1:00:57] I get very excited, read all the things about it, write about it, and that's mainly it. I almost go by intuition, basically. And I'm kind of lucky in that sense that with my blog, what I find interesting, other people also right now find interesting. So I think there's a lucky coincidence that I honestly write only about things I find interesting. So I'm not trying to force myself, oh, I have to cover XYZ because it's something that should be covered.

Sebastian Raschka [1:01:25] It's more like, oh, how does this, let's say, recursive language model work? Let's just read the paper and then I write about it. More like getting excited about things. Yeah, the book writing is also a bit different because my blog is more research-paper-focused, where when I get excited about something, I read about it and put it in there. For the book, I'm similarly excited, but that's more like a coding book where it's like the fundamentals.

Sebastian Raschka [1:01:54] Because I think that's, to be honest, the best way to understand something: to see it actually working. It's not any hand-wavy figures. I mean, there are a lot of figures. Like right now for chapter six, I just finished it the other day, 21 figures. They take the most work. Maybe one day LLMs can help me with that. But figures help to explain the code and everything. But code basically doesn't lie. It either works or it doesn't work, and I think that's a very useful way to learn also.

Sebastian Raschka [1:02:19] And it's just also a lot of fun for me. When I write code and I see it working, it's very satisfying, and you have something that actually works. So I should say I'm not building, in that book, any production-level systems. It's called Build an LLM from Scratch, or build, I mean, large language model from scratch. But it's not like an LLM that you would use in production. Well, if you make a few tweaks, you could use it in production.

Sebastian Raschka [1:02:46] But the goal is code readability and teaching LLMs, basically, because I think to see actually, how do I format my training data? How does it get processed? What is the loss function? What gets updated? I think this explains so much more than if I say, oh, so it does next-token prediction, and then hand-wavy, hand-wavy here and there. You can actually literally see how it does that and what feeds in. And the same with RLVR. We had hopefully not too bad an explanation in this podcast at the beginning of GRPO.

Sebastian Raschka [1:03:13] But if you actually see the diagram, I have the numbers in there, but the numbers could be wrong, though, if you have a figure and you draw arrows and everything. But then if you implement that in code and you get exactly the same results and the model trains and you get 50% accuracy, it's a nice thing where you can say, oh, it's actually working. It's not just made-up numbers. It actually works. And also that's how I learn about LLM architecture.

Sebastian Raschka [1:03:38] So I have this blog, The Big LLM Architecture Comparison, with now, I think, 13,000 words because I keep extending it. I read the paper, I look at the architecture of an LLM. Do I really understand it? So I draw the architecture. But then do I really understand it? And often, unless it's like a 1 trillion parameter model, I often code the model. And so I have still my GPT-2 architecture, and they are relatively similar. And so if there's a new architecture, I take the most similar one and make a few changes to it.

Sebastian Raschka [1:04:07] But then the beautiful thing here is also someone already implemented that in Hugging Face, the Transformers architecture. So I have a reference model I can run. I run my model and I can see, do I get the exact same logits, the same numbers, if I have the same prompt? And so with that, you can actually self-check yourself. Did I implement everything correctly? Are the results correct? And I think that's just a lot of fun. It's a lot of work, but it's a lot of fun.

Sebastian Raschka [1:04:31] And it doesn't lie. It gives you the correct answer. Perfect. On some prompts, someone else extended that and I found a bug. And now I have a better understanding of how they implemented the YaRN scaling. That is something you would never understand by just reading the paper. You have to really, I don't know, look at the code and toy around with that. And so, yeah, that's basically how I try to work. I try to combine reading and coding, and LLMs I also use, but I try to use it.

LLM burnout + how he uses models (without replacing himself)

Sebastian Raschka [1:04:56] So I would say for blog writing and book writing, not so much. Because honestly, for fun, I tried it out. It generates okay text, but I don't know, I can ask it to generate text like me, but then I don't like it, and I end up editing it. It's almost faster for me to just write it out the way I want it.

Matt Turck [1:05:02] And you had interesting thoughts on LLM burnout, how using LLMs tends to deplete energy.

Sebastian Raschka [1:05:26] I would say first, also, the thing is, it's not super satisfying if you just ask the LLM to do it. It's like cheating at homework. Honestly, I understand there are jobs and people where it just matters how much you get done and how fast you get something done. And then it makes sense to use an LLM to do the job for you. But I think there are different types of people who enjoy different things. For example, I enjoy doing the work more than managing.

Sebastian Raschka [1:05:51] When I was a professor, I did research, but I also had to supervise other students. And I liked working with other students, but I noticed I actually liked doing the research myself more than telling other people how to do research or managing. And so I think if you use only LLMs to just generate everything, I wouldn't say useless, but I would feel maybe empty. Like, okay, you lose that pride, I think—the pride of, oh, I did something that worked, and it's cool, and you're proud of this.

Sebastian Raschka [1:06:19] And so what I try to do is generally, when I use LLMs, I try to make my work better, not necessarily to make more or make it faster. I mean, to some extent I do. But then I try to think, how can I make what I do kind of better? So what I use LLMs for is more like, hey, I use actually GPT-5 Pro when I've written an article and put it in there. Hey, can you find any mistakes or typos?

Sebastian Raschka [1:06:41] Often I have misnumbered, mislabeled figures. I go, like, figure 11, 12, 15, 16, and things like that. I can check myself, but it's just faster for an LLM to find all these things, like how to make things better, or are there any things that are unclear? I mean, I'm not a native speaker, so sometimes I have problems with a sentence. I'm tired, I just can't get it right, and then it suggests to me, oh yeah, that is maybe not a bad way to say it, and I would then take that sentence, for example.

Sebastian Raschka [1:07:17] Things like that, where I'm trying to make, let's say, the work better without fully replacing myself. I mean, maybe that's shortsighted because LLMs will eventually be able to do everything, but I kind of enjoy the work too much to just give everything away, if that makes sense. And I have the luxury that this still works for me. So I know there are some businesses where it is really important to execute faster.

Matt Turck [1:07:41] Thank you very much for your work. You're very popular for the quality of your writing, all your tweets. You have a big X following, and your name keeps coming back in conversations about where people go to learn. Thank you for doing this part. Really appreciate your time. That was super fascinating and really educational, really insightful. So really appreciate it. Thank you so much, Sebastian.

Sebastian Raschka [1:07:50] Thank you for inviting me. I had a lot of fun. Five hours just talking about LLMs and AI. I mean, that's the dream, right? Thanks for having me.

Matt Turck [1:08:13] Thank you. Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.

Sebastian Raschka [1:08:13] Bye.