The End of GPU Scaling? Compute & The Agent Era — Tim Dettmers (Ai2) & Dan Fu (Together AI)
The MAD Podcast with Matt Turck · with Tim Dettmers, Assistant Professor, CMU; Research Scientist, Allen Institute for AI
Tim Dettmers is the Assistant Professor, CMU; Research Scientist, Allen Institute for AI. We cover why current models may leave up to 100x more training compute available through newer clusters and better utilization, how coding agents make kernel engineers up to 10x faster with expert oversight, and why inference hardware utilization below 5% leaves major room for cheaper, faster models.
Chapters
- 1:06 — Two essays, two frameworks on AGI
- 1:34 — Tim’s background: quantization, QLoRA, efficient deep learning
- 2:25 — Dan’s background: FlashAttention, kernels, alternative architectures
- 3:38 — Defining AGI: what does it mean in practice?
- 8:20 — Tim’s case: computation is physical, diminishing returns, memory movement
- 11:29 — “GPUs won’t improve meaningfully”: the core claim and why
- 16:16 — Dan’s response: utilization headroom (MFU) + “models are lagging indicators”
- 22:50 — Pre-training vs post-training (and why product feedback matters)
- 25:30 — Convergence: usefulness + diffusion (where impact actually comes from)
- 29:50 — Multi-hardware future: NVIDIA, AMD, TPUs, Cerebras, inference chips
- 32:16 — Agents: did the “switch flip” yet?
- 33:19 — Dan: agents crossed the threshold (kernels as the “final boss”)
- 34:51 — Tim: “use agents or be left behind” + beyond coding
- 36:58 — “90% of code and text should be written by agents” (how to do it responsibly)
- 39:11 — Practical automation for non-coders: what to build and how to start
- 43:52 — Dan: managing agents like junior teammates (tools, guardrails, leverage)
- 48:14 — Education and training: learning in an agent world
- 52:44 — What Tim is building next (open-source coding agent; private repo specialization)
- 54:44 — What Dan is building next (inference efficiency, cost, performance)
- 55:58 — Mega-kernels + Together Atlas (speculative decoding + adaptive speedups)
- 58:19 — Predictions for 2026: small models, open-source, hardware, modalities
- 1:02:02 — Beyond transformers: state-space and architecture diversity
- 1:03:34 — Wrap
Transcript
Two essays, two frameworks on AGI
Matt Turck [1:05] Tim and Dan, welcome.
Tim Dettmers [1:15] Thanks for having us. So Tim, a few weeks ago, you wrote a great, provocative blog post entitled "Why AGI Will Not Happen."
Matt Turck [1:23] And then Dan, a few days later, you replied with your own blog post, equally fascinating, entitled "Yes, AGI Will Happen."
Tim Dettmers [1:26] I'd love to go into your backgrounds.
Tim’s background: quantization, QLoRA, efficient deep learning
Matt Turck [1:58] You both have the very interesting characteristic of having a foot in industry and a foot in academia. Tim, if you want to start with yours. I'm an assistant professor at Carnegie Mellon University in the Machine Learning and Computer Science departments, and also a research scientist at the Allen Institute for AI. My past research has been mostly on efficient deep learning quantization. That means model compression: taking large models and compressing them down from 16-bit to something like 4-bit. My key research has been there. QLoRA, for example, is very efficient fine-tuning compressed to 4-bit.
Dan’s background: FlashAttention, kernels, alternative architectures
Matt Turck [2:26] Use adapters on the model and then use up to 16 times less memory than if you have dense fine-tuning. And now I'm working on coding agents. There, we have a very exciting release in about two weeks: state-of-the-art agents you can quickly specialize to private data, get strong performance on any codebase that you like. And yeah, that's very exciting. Great. Dan?
Tim Dettmers [2:50] Hey, so I'm an assistant professor at UC San Diego, and also my title is VP of Kernels at Together AI. In industry, I focus a lot on basically making models go fast. GPU kernels are the things that actually translate the models to how they run on the GPU. You can think of them as basically specialized GPU programs. A lot of my research in my PhD and in my lab focused on that. I developed things like FlashAttention, which was an efficient kernel for one of the core operations of a lot of the language models that we use today.
Tim Dettmers [3:19] I also did research on alternative architectures to transformers, things like state-space models and things like that. At Together, I'm really focused on how do you make the best language models that we have today go faster on NVIDIA's Blackwell GPUs. That's a bit of a flavor of what I do.
Defining AGI: what does it mean in practice?
Matt Turck [3:55] Let's get into this AGI discussion, and then in the second part of this conversation, we'll talk about agents and coding agents and your thoughts there, because I want to make sure we cover that. AGI, obviously, is a term that everybody uses, and I think we can all agree that nobody really knows what that means. But for purposes of this discussion, what is a useful definition of AGI from your perspective?
Tim Dettmers [4:17] Sure. Yeah, I think one of the things that we kind of discussed back and forth in this set of blog posts is sort of what AGI means. For me, I think one of the things that I've been thinking about recently is that if you took where we are today with the models that we have today, with the language models—and I think we'll probably talk about this a bit more later with the agents—by almost any definition anyone could have written down, let's say five years ago or 10 years ago, certainly when Tim, you, and I started our PhDs, we basically have the vision of AGI that we had back then.
Tim Dettmers [4:54] We have things that can write code. They can write human text, even though maybe they use too many em dashes or something like that, but they can do these really amazing things. I think one of the things that I think about is: at what level does this become a new industrial revolution, where this technology is really going to change a lot of the way that we do things today and have a huge, really great economic impact? In terms of software engineering, I feel like we're already there or almost there.
Tim Dettmers [5:21] There are things that may be super specialized. I don't know if they're going to be able to write the best Fortran and COBOL code in the world. But for web development, even a lot of low-level systems engineering, they're already really great. One of the reasons that I wrote my blog post was: if you think about where we are today, we maybe already have AGI or some form of AGI. If not, then certainly the next generation of models, the models that are training today, if they're at all better than what we have today, then we've already hit something that's really amazing and pretty wild.
Matt Turck [6:01] When I wrote my blog post, actually, I noticed, like, oh, I forgot to put the definition of AGI in my blog post, even though my blog post is very much about AGI. And I think that sometimes sort of reflects how we think about AGI. We don't think carefully about the definition. I mean, there are sort of—and I thought about it before—and I think there's sort of different kinds of definitions that have advantages and disadvantages. I wouldn't say, as you said before, there's not one definition that people agree on, just to mention a couple.
Matt Turck [6:31] And I think one that's sort of quite widespread is to see AGI as cognitive abilities, cognitive tasks. What can you do cognitively? And software engineering, very cognitive; writing, very cognitive; moving a robot in space, that's more kinetic. You could also say, like, hey, you also need to think about how you move. That's also part of cognition, but I think most people would separate that and say everything digital is kind of cognitive.
Matt Turck [7:08] And if you have physical, that goes beyond that. What I think makes sense is sort of this economic angle. Can we get another industrial revolution? What it means is: is AI useful? And is it so broadly useful that you want to use it everywhere, kind of, and it accelerates all kinds of things? Similar to when computers were introduced, productivity increased. Not initially—productivity actually went down—and you need diffusion in the economy to pick it up again. We might see something like that with AGI more broadly, and tasks like software engineering scaling up pretty significantly.
Matt Turck [7:21] But yeah, I think that is useful. Let's jump into the heart of the argument, Tim.
Tim Dettmers [7:31] I was amused by what you said about where all those ideas of AGI and superintelligence come from, if you want to talk to that.
Matt Turck [7:54] Yeah. And to sort of lay out the entire narrative, there are certain thoughts about AGI, and that is sort of very much rooted in a certain kind of thinking. It comes from effective altruism communities and the rationality communities. I was part of these communities a long time ago; that is now 15 years ago. If you look on Twitter, there's always like, oh, we get AGI in two years. And then one year later, oh, we get AGI in two years.
Tim’s case: computation is physical, diminishing returns, memory movement
Matt Turck [8:21] And then one year later, we get AGI in two years. I feel like it's a little bit lazy thinking, a little bit of being in a bubble and not being exposed to different ideas. And that was one of the main motivations for me writing the blog post, because I feel like there are some ideas that, if you think about them, might provide a counterpoint to a lot of the thinking that's out there. Yeah. And your core thinking is that there is a tension between those ideas and the computational reality.
Tim Dettmers [8:29] Is that a fair way to put it?
Matt Turck [9:00] Yeah. There's a physical component, and then there's an idea component, but they have a very similar structure. And this structure is basically diminishing returns. Everything that grows exponentially will level off because, if you need resources, the resources will be exhausted. Resources can mean different things. And if you look at the physical aspect, it gets more and more difficult to advance technology. That is the case almost within any field of research or development. Things get more and more difficult; you need more resources to make further progress, and the progress sort of goes lower and lower.
Matt Turck [9:42] And so, if you look at the physical reality of computational devices, and then also computation itself has a particular structure. And so basically, useful computation is two things. The first is, you need to gather data from one location and aggregate it in a certain location, where you then put this new information together to compute a transformation of that information. You basically want to combine known things and compute new things that you didn't know before: useful information.
Matt Turck [10:15] Useful information can and needs to be transformed from information that you already know. If you move a lot of information around, but you don't transform it, you can't make new information. If you do a lot of computation on the information that you already have, you miss out on the long-distance insight, the indirect insight. I think a lot of this actually maps to the neural network architectures that we have. In the beginning, we had convolutional networks. They're very effective. And what they're doing, they don't move much memory.
Matt Turck [10:47] They do a lot of computation. And that means your device needs a lot of FLOPS, and memory bandwidth is not that important. Once you go to very dense computation, very large matrices, then it goes in the direction of recurrent neural networks. But there, you still have this component of your recurrent network basically paying attention to previous states. But because it's recurrence, the memory reuse of that computation is minimal. And with transformers, you basically then had these large matrices that compute, basically, that transform the incoming information from the previous layer.
Matt Turck [11:27] And then you had attention that now computes information across time or space. And I would argue these are the two most fundamental ways of computing information. You want to relate information to itself, or transform that information, but then you also want to basically relate information to distantly related information. So you want long-term relationships, and you want to have transformations based on what you already know.
“GPUs won’t improve meaningfully”: the core claim and why
Tim Dettmers [11:29] And you say this is slowing down, right?
Matt Turck [11:37] Like in your blog post, you have a pretty striking sentence where you say GPUs will no longer improve meaningfully.
Tim Dettmers [11:40] We have essentially seen the last generation of significant GPU improvements.
Matt Turck [12:04] Yeah. So this has two components. One is also sort of a very fundamental thing, and it's physical, in the sense that I mentioned these two components: memory movement and computation. Computation can only be useful if you move memory to this sort of local neighborhood where you do this computation. Now, this is a geometric problem. You need to have a large store of information and then use this large store to move information closer to where you do want to do the computation.
Matt Turck [12:44] And we have figured out how to physically do this optimally. We have a large, slow memory—that's DRAM. Then we move it to a cache. If you look at the geometry, that's how you do it fast. If you have a certain size of computation, this is optimal. If you have a different size of computation, matrix multiplication, then you want to use not a CPU but more like a GPU, which has higher latency but more throughput. You can move more data, but more slowly.
Matt Turck [13:17] And if you look at all of that, you can push around a little bit how you structure everything, like the caches and how large they are and how many cores they are shared across. But in the end, the fundamental problem remains the same. You have a geometric problem. You can only fill the space in a certain way. And that means you always have certain access patterns with certain latencies. And the biggest latency is a big block of DRAM. That is the major bottleneck.
Matt Turck [13:47] This is also called the von Neumann bottleneck, based on computers—almost all computers that we have. And this is the bottleneck of moving a program to where you execute the program. And for neural networks, that's basically moving the weights and the inputs to the execution where you execute the program. That will be the tensor cores. There are not many ways you can go around this bottleneck. The only way is to store the memory locally and do the local computation there.
Matt Turck [14:08] And there are some processors that do something like that. For example, this Cerebras processor. So they don't have this von Neumann bottleneck in a major way on the chip, but then they need to also pipe data into that chip.
Tim Dettmers [14:08] Right.
Matt Turck [14:36] And so the von Neumann bottleneck moves basically away from the chip to your storage or to your network. And so you just move it away, but it's still the same bottleneck. You need to load the program, which might reside on disk or memory, through the network to the chip. Same physical problem. You just move a couple of variables around. That is sort of one part of the problem. We don't have architectures that can solve this problem. That's sort of the second part where my argument kicks in.
Matt Turck [15:07] And that is, you need new technology to overcome bottlenecks. But once you have leveraged that technology, you need new technology to get over that. If you look at what we can do, we moved from DRAM to HBM. So that's DRAM that's stacked. That's much faster, but you can only stack it that high because it's very difficult to manufacture and test for correctness. And yield is very low. And if you run, actually, in 2026, there's not enough HBM. You can't build the nice processors anymore because you run out.
Matt Turck [15:41] It's just too difficult to manufacture them. And with that, we have all these innovations. One of them was Tensor Cores, a big step up. Then we have 8-bit precision, another step up. Then we have 4-bit precision with particular block-wide quantization, particular data types. From my research and other research, we know that's close to information-theoretically optimal in a practical sense. If you train on enough data, 4-bit precision is not enough. You need actually 8-bit precision. So you can't go further.
Matt Turck [16:09] The hardware is maxed out. We have no new technology. We can make it easier to manufacture, a little bit cheaper, but not faster. And you have maxed out on the additional features. Sparsity could be something. People tried it for 50 years. I tried it. It doesn't work. And so that might be the last thing, but 4-bit precision is the end of quantization. And so that's the end of it. We have maxed out the features. We have maxed out the hardware.
Matt Turck [16:13] That's what we get. Okay. Fascinating.
Dan’s response: utilization headroom (MFU) + “models are lagging indicators”
Tim Dettmers [16:35] All right, Dan, what is your perspective on all of this? I really appreciated Tim's post because I think one thing that I really appreciated is that there's some AGI talk that, if you just trace the exponential, at some point you get the thing that will eat up the universe or whatever, which I always found a little bit odd to think that way. I appreciate the thing in terms of the actual physical constraints because, like Tim said, these are physical systems with physical inputs and actually doing physical computation.
Tim Dettmers [17:10] I think my perspective was that if you look at where the systems are today and you look at the models that we've trained, we are just so far from even using the last generation of hardware as efficiently as possible, not to mention all the new hardware that's being built out. On the technical side, I'd say there are two major points I wanted to make in my post. One, if you look at the models that are the really great ones that we know today.
Tim Dettmers [17:44] In my blog post, I mostly talked about open-source models because they talk a little bit more about how they train and the resources behind it. We don't have public figures behind how much OpenAI and Anthropic are using. But if you look at the DeepSeek model, for instance, this is one of the best open-source models we have out there today. It was trained at the end of 2024 on last-generation, kind of nerfed GPUs, H800s instead of H100s. The H800 is nerfed in all sorts of ways from NVIDIA to get around the export restrictions at the time.
Tim Dettmers [18:15] They were trained with about 2,000 H800s, according to the report, for about a month. When you compute how long that took, when you see how much compute was actually available on the chip, you get something like a 20% effective chip utilization or something like that. The term of art is called MFU, model FLOP utilization. Basically, that's a 20% utilization number. Meanwhile, earlier in the 2020s, we were seeing lots of training runs on older hardware with different model architectures that were easily achieving 50–60% MFU.
Tim Dettmers [18:42] If you just take that number and then say, hey, maybe there's a way to get it out there. Since then, my good friend Tri Dao has released a whole new set of kernels on how to train these models better. And you say, okay, there's a 3x there just from that one piece. Then the other thing to realize is that that is a model that is being used today in early 2026 as the base for some of the best open-source or open-source-adjacent models out there.
Tim Dettmers [19:21] It would have started training the base model at least a year and a half ago. So let's call it mid-2024. Since then, we've started building out completely new clusters with the current generation of hardware. On NVIDIA, these are the Blackwells. There are companies like Poolside that are building out tens of thousands of B200, GB200 chips. There are other folks like Reflection who are building out tens of thousands of B200 chips. This is comparing: we have a new generation of hardware where, even if you take the exact same precision as you had before, exact same everything, 2 to 3x faster compute, 10x larger clusters, plus maybe 3x lurking in terms of just pure optimization.
Tim Dettmers [19:56] That's 3 times 3 times maybe another 10. That's another 90x of compute available. And that's not even looking at future buildouts. That is literally clusters that you can point to today that people have started training on that you might hope that at the end of that you'll get much better models. The point I really wanted to make was: if you just look at it from those basic inputs, you can look around, you can squint a little bit, you can see up to two orders of magnitude more compute available compared to the models that we are indexing on today.
Tim Dettmers [20:39] Now we can argue about, are there going to be diminishing returns in terms of scaling up? Do we expect the scaling curves to hold and all that? But you can just look around and see it. That's 100x more compute. I think from the physical, just a pure compute perspective, there's a lot more available, a lot more that we're not doing. This is not even to mention a bunch of the great points that you mentioned, Tim. These are all 8-bit training runs.
Tim Dettmers [21:11] We've just started writing the papers about how to do a 4-bit training run properly. There's new things like, on the GB200, you have 72 of these connected really quickly. I don't think we've even seen the first pre-trained model come out of that yet. GPT-5, I think, was the first time that you saw in one of OpenAI's reports, hey, this was trained on H100, H200, and GB200, which to me suggests that was actually pre-trained on one of the really old clusters. Maybe some fine-tuning was done on the new GB200s.
Tim Dettmers [21:23] You make the point that not only is the hardware underutilized, but you also say that the models themselves are a lagging indicator. The models that we see today that we can play with today have been pre-trained on clusters that were built out a year or two ago because you need enough time to get the cluster running, you need enough time to do the large pre-training run, and then you need enough time to really post-train it, fine-tune it, do all the RLHF and all that stuff.
Tim Dettmers [22:04] So the models that we have a snapshot of today that, at the beginning of a conversation, say maybe it is AGI, maybe it isn't, are already trained on clusters that are a year and a half old. We've built out much larger clusters since then. You can expect that they're going to use them for pre-training. The models that we see today, that we index on quality today, are actually trained on pretty old hardware, and we've got new generations of hardware, more software choices we can make, not to mention architectural choices.
Tim Dettmers [22:45] Tim, you were mentioning this thing about, you need to move data and then compute on data. We've actually seen the transformer change in architecture a little bit slowly for researchers, a little bit slowly for my taste, but you've seen the fundamental way we do the computation change. Five x or two x there, now you're talking 100x, 150x more compute. There's a lot more compute out there to train better, higher-quality models. If I understand this whole discussion correctly, all of this is about pre-training, right?
Pre-training vs post-training (and why product feedback matters)
Matt Turck [23:09] And whether we can train a bigger model with more data and more compute. But in conversations on this pod, a lot of the conversations have been about the importance of post-training and building AI systems with pre-training plus RL.
Tim Dettmers [23:36] Where does that fit? That's a great question. And I think another piece that I don't think either of our blogs particularly hit on: one way I like to think about it is that pre-training is like the general strength training that you do in the gym. You go lift heavy weights, you improve your strength, improve your general ability. And then post-training is like the specific drills that you run to get good at a particular task. Historically, the vast amount of compute has gone to pre-training.
Tim Dettmers [24:02] Just building models that are more generally capable of doing many things, have a lot of knowledge, get to a point where maybe they have more knowledge than your average person. I certainly don't know as much as ChatGPT, for instance. Then post-training is also: how do you make it helpful? ChatGPT, you ask it to do something, and then it actually listens to you and tries its best to do it. But I think the other thing that we've started to see increasingly in post-training is that you can start to post-train specific skills.
Tim Dettmers [24:38] The model that's really good at helping you code uses a lot of the knowledge that you got from pre-training, but is actually adapted to be particularly good for coding. Or the model that's really good for legal work, for instance, has a lot of the pre-training backbone, but then the post-training is really what gets it to that place where it's really useful. From a pure computational perspective, pre-training is usually much more compute-intensive than post-training. Post-training, the work that you have to do, I think—I'm not a post-training expert—but the work ends up looking a lot more like: how do you build a useful product?
Tim Dettmers [25:13] How do you get user feedback? How do you do things like that? Even then, there's a world where maybe the next generation of pre-trained models is a strong enough base that if you go tackle each vertical of the economy that you care about, you could actually post-train it to something quite useful. I think that's a whole other computational aspect of it. Maybe we don't even need that 100x more compute that may be out there. Maybe it's more traditional work of, let's understand this problem deeply.
Convergence: usefulness + diffusion (where impact actually comes from)
Tim Dettmers [25:45] Let's understand how to train in almost the human sense. How would you take an intern and train this intern to do this specific task? How do you get this very powerful pre-trained model to do something really useful in this post-training sense? It's this concept of usefulness that you both mentioned where both of your points of view maybe converge. In some ways, AGI is something, but what ultimately matters is where you land in terms of usefulness in the industry. And therefore, even though one may not be able to reach that kind of ethereal definition of AGI that nobody really understands due to diminishing returns, in some ways it doesn't matter because we still have so much juice to squeeze that we have enough to go until we get to a place where this is truly useful, not just for coding, but for the rest of the economy.
Matt Turck [26:31] Yeah, the main conclusion of sort of my blog post was exactly that. You shouldn't pay too much attention to AGI, but more about thinking about how we can make it most useful. That might go beyond how useful a model is. I mean, Dan mentioned post-training as a product. An important part we saw with computers is diffusion in the economy. That requires a very different mindset. The U.S. mindset is build the best model, and then everybody will use it.
Matt Turck [26:54] But that can help to really figure out: how can you benefit the most people in the most pragmatic way? And I think that's sort of more of a Chinese mindset. And so, that sort of mindset. So if I think of usefulness, one is model, the other is sort of the mindset. But I would agree. I think that both Dan and I, and I think most people, would agree that if you have AI that does very impressive things like Math Olympiad things and that sort of thing, but it can't do anything useful, is it AGI?
Matt Turck [27:25] And so models are already useful, so that scenario will not happen. But I think what we really want is very, very useful models. And I think we have that, and I think we can prove that. But I don't think we get to AGI by certain definitions, but we will see significant impact.
Tim Dettmers [27:48] Yeah, I think I would just add to that, that Tim, you had this point about how much of the economy is physical and how much of it is knowledge work. I think that the U.S.-China contrast is really interesting there. There's been these analyses, this book by Dan Wang going around, about the manufacturing economy, the engineering economy versus the more lawyerly economy. I think there's certainly a lot of great knowledge work to be done in the U.S. I think there's also, if you look at what the actual sectors of the economy are, a large portion of it is healthcare, a large portion of it is education.
Tim Dettmers [28:29] Tech is certainly also a large portion that's leading the stock market and driving the stock market. There's a lot of great people who are trying to use the new models to try to do things like develop new drugs or understand how to make a real impact in healthcare, or if we can get robotics off the ground and do things like start helping with some of the physical labor, maybe not necessarily building houses, but the day-to-day household labor. That could be large, untapped portions of the economy.
Tim Dettmers [28:55] Those pieces are really great. You can almost start to see the first pieces towards it. The self-driving analogy is really interesting to me because early on in my PhD, I was quite skeptical about self-driving. So let's call this 2018, 2019. It felt like self-driving was always a year away or two years away, or if you ask the experts, they'd say, oh, five years away. And then last year, I rode in a Waymo, and today I just actually got access to Waymo on the highway.
Tim Dettmers [29:18] So now conceivably, I could potentially sell my car, and I live in the Bay Area in California. I won't because I personally like driving, but the progress is funny in this way where it's kind of, it's not there, it's not there, it's not there. And then one day, a switch flips, and then you're suddenly like, oh, not only is this thing pretty good, it's actually a lot better than the service that I'd get in an Uber or a taxi or something like that.
Multi-hardware future: NVIDIA, AMD, TPUs, Cerebras, inference chips
Tim Dettmers [29:59] That's a really exciting thing. If we see that happen, if we see that switch flip for household cleaning or putting away the dishes or things like that, I think it would be really exciting. It would change a lot of folks' perspectives. I'm not a roboticist myself, but I'm really watching that space with a lot of excitement. Dan, as a quick tangent, do you think we're evolving towards a multi-hardware, multi-chip kind of world based on what you see? Obviously, there's Groq and NVIDIA, there's Cerebras, there's a bunch of sort of special ASIC companies coming up.
Tim Dettmers [30:32] From your kind of low-level-in-the-stack vantage point, what do you see? Yeah, that's a great question. It's something that I spend quite a bit of time thinking about, more so, I'd say, on the lab side than necessarily on the industry side, although, of course, we're paying close attention on both sides. I think it's a really exciting time where the NVIDIA chips are really strong, really reliable. There's a lot of software support around them that has built around them.
Tim Dettmers [30:58] We're starting to see the same things happen, for example, on AMD chips with some of the research there. On the lab side, we recently put out a library called HIPKittens, led by my great friend Simran Arora. She was really looking at what are the right software abstractions to program on these AMD GPUs. It turns out they're not exactly the same as the NVIDIA GPUs, even two GPUs that have relatively similar specs, certainly compared to Groq or Cerebras or SambaNova or one of these other chips.
Tim Dettmers [31:35] Even though they're relatively similar, they actually have pretty different software abstractions you need to use. I think more people are getting excited by that and investing time and energy into that. We saw the Groq acquisition from NVIDIA. A lot of people are excited about TPUs today. I think Cerebras and OpenAI just announced their partnership. So I think certainly it's going to be a wave of things coming forward that you're going to see a lot more. I'm sure NVIDIA will still do great and still grow beyond their $5 trillion company, or whatever it is at the time of recording.
Tim Dettmers [32:04] But I think you're going to see a lot more diversity, especially around inference of the model. Training and inference are actually quite different computations, and as a result, you might actually want quite different chips to do it. On the inference side, you might want, for example, your models to live locally on your phone, on your laptop. My phone, my iPhone, which is a few years old at this point, is already more powerful than some of the GPUs that I had when I was starting my PhD.
Agents: did the “switch flip” yet?
Tim Dettmers [32:24] That growth of that hardware power is really exciting to see. Dan, you mentioned a second ago, in reference to self-driving cars, that moment where things flip, the switch is turned on.
Matt Turck [32:28] Has that happened with agents already?
Tim Dettmers [32:48] You talked about software singularity. Are we at that moment for agents? Yeah, I think that, so personally, in my life, I'd say that moment was last June-ish. So June 2025 was the moment that it really flipped for me. To give some context here, what I do in my day job at Together AI is we write a lot of these GPU kernels. I don't know how popular, but in the general ML zeitgeist, GPU kernels are thought of as kind of like the final boss of the thing that you learn how to program.
Dan: agents crossed the threshold (kernels as the “final boss”)
Tim Dettmers [33:21] They're very hard. They're very highly parallel. You have to write them in C++, which is this old language that the old systems people used decades ago or whatever. They're not in Python, et cetera. When you're trying to hire people who can write kernels, it's very hard. It's a very challenging skill set. It's certainly the tip of the spear in terms of programming strength. Last June, we had this really interesting realization where we realized that Claude Code, Cursor Agent, these agentic coding assistants, were actually very good at writing these kernels.
Tim Dettmers [33:49] There was one week where I think I wrote three or four different features that usually would have taken me a week each. I wrote all of those in a single day. I was like, oh my God, this thing is making me five times more productive as a kernel expert. I got my team on it. Now my team has all these really complex systems that they've built where they can write a whole feature that I think would have taken months of a whole team's time before.
Tim Dettmers [34:28] This is kind of that final boss of programming challenge that was really challenging. From our perspective, for coding, for this really technically challenging GPU kernel programming, it kind of crossed the Rubicon for us already. I gave this talk a few months ago at Slush about what we're calling the software singularity, where we realized, hey, in terms of software engineering, even for these really niche skills, it's certainly better than the average programmer. It's at a place where it can accelerate the really expert programmers.
Tim: “use agents or be left behind” + beyond coding
Tim Dettmers [34:51] Right now, as of today's recording, it's at a place where if I just let it on its own, it might not generate the right thing for you. But if you give an expert programmer this set of tools, they can go 10 times faster than they were able to go before. And I think that that's a really exciting place to be.
Matt Turck [35:07] And on that topic of agents, Tim, you just wrote another great blog post called "Use Agents or Be Left Behind." And part of what you talk about is coding agents versus agents for the rest of our lives.
Tim Dettmers [35:16] Where are we in that arc? Agents are transitioning from being excellent at code to useful for the rest of our lives.
Matt Turck [35:38] That blog post was also a reaction to what I see: there are a lot of productivity gains if you use coding agents for all kinds of tasks. And as a professor, you don't code that much. You can actually code more easily, which probably other professors previously would not do. It's so easy now. But yeah, also for non-coding tasks, it's super useful. And when I look at the productivity gains that I have, some of it's smaller, like two or three, sometimes it's like 10 times faster.
Matt Turck [36:13] I do tasks 10 times faster. The quality is not degraded. Sometimes the quality is higher. An agent might not be as good as I am, but the agent doesn't get tired. The agent doesn't make sort of bad mistakes or need to cognitively struggle with complicated information that you put together, similar to CUDA kernels, what Dan mentioned. All of that is working. And I mean, Matt, as you put it, it's coding agents and agents for other stuff. But how I would see it is, it's just coding agents.
Matt Turck [36:43] Coding agents are general agents. Coding agents can write programs that solve other problems. And code is so general. If there's a digital problem, you could solve it with code. And coding agents make things so easy that now you can solve a variety of problems in a way that you couldn't solve before. And this angle makes you productive. I would say this is the main thing that has been changing. Coding agents allow you to attack a problem in a way that you couldn't think before.
“90% of code and text should be written by agents” (how to do it responsibly)
Matt Turck [36:59] And it's at a pace that you couldn't think before. You can parallelize a lot of different tasks. The agent doesn't get tired. You just keep going. The work's much easier.
Tim Dettmers [37:19] One bit I love in your post is that you're careful to separate the hype from reality at the beginning, but then quickly you land, from your experience of experimenting with agents for the livestream in particular, at the conclusion that more than 90% of code and text should be written by agents. You need to do so or you will be left behind.
Matt Turck [37:44] I think for many people who are engineers, that's already true. And there is this thinking: if you produce text or code and everything's done by agents, it must be low quality, must be bad. But the key thing is you inspect the code, you inspect the text, you might make some slight edits. The 10% that you do might make a big, big difference through this basically sort of editing, just reviewing of output. You kind of make it your own.
Matt Turck [38:07] AI-generated things are not less personal than things you've written on your own. I see that. If I write a grant proposal with the help of an agent, it's sort of alive. I can feel like it's exciting. The person reading it says, like, this is good research. I want to fund this. I think that's just the reality. If you just generate things and don't look at it and just say, like, yep, that's good, that will not help you.
Matt Turck [38:41] But you can quickly review content, you can skim it, you can look at, like, ah, this doesn't look right, or I want it different, and edit it, and you're good to go. That will be the reality, and the skills that you need to work in that way, they're not fully developed for most people. They're also not fully developed for me. It's still sort of a phase of experimentation. Models move, frameworks move, and so you need to adapt. You need to learn.
Matt Turck [39:09] There's a lot to learn, but if you do it, the payoffs are huge. And I think there was a thinking that software engineers will no longer exist, but I think people no longer believe that. It's like software engineers are so productive. That is exactly what you need to learn. If you use agents well, you can do so many things. I think that's the core thing. If you don't know how to use agents well, you will be left behind. That will become a critical skill.
Practical automation for non-coders: what to build and how to start
Tim Dettmers [39:25] Practically, how do I do that if I'm not a coder and I think about automating some parts of my job? What are some of your recommendations on how to approach that problem?
Matt Turck [39:48] Yeah, I mean, the best thing is just being very pragmatic. Just think about things and try to code them. Particularly if you're not a coder, that's very difficult. And there's this barrier where you say, like, I haven't coded before. I don't know this. But if you interact with agents, they can just build stuff. And with minimal learning, they can also explain stuff. With minimal learning, you can get there, execute programs, build websites. Particularly if it's visual, you get quick feedback.
Matt Turck [40:19] It's not that difficult anymore. I mean, often I mention you need to inspect things, but if you build simple tools for yourself to make your life easier, often you don't need to do that. The agents write good code. If you work in a company, you need to integrate it into a good codebase, you probably should review it. But if you build a small program on your own to make your work more productive, that's easy. Just to give an example that might be relevant here, I built a tool that, if I have a video where I talk—I record videos of how I interact with agents—there are certain phases where I just look at outputs and try to understand things.
Matt Turck [40:56] There are phases where I talk. So I just built a tool that recognizes the speech, determines the timestamps when I'm speaking, then slices the video. So basically, I have an entire video where I talk rather than moments where nothing happens. And that's very easy to do. I built this in 20 minutes. I think everybody can do it because I didn't look at the code. The agent did it. And then I look at the video, and I'm like, oh yeah, it did it right.
Matt Turck [41:14] If you get started with a feedback loop, you don't need to code. You just need to inspect the output that you can understand, or learn how to execute a Python program or a Bash shell, and you're there. How do you pick what you want to automate?
Tim Dettmers [41:17] How do I think about automation in my life?
Matt Turck [41:37] Yeah. So I also talked about that in a blog post, and it can be a more intuitive thing and then a more nuanced thing. I think the more intuitive thing is, like, you just think about what could be useful. And then it can even be something more complex. You say, I want an Android app or an iPhone app that does this thing. And you initially might think that's complex, but then you throw it at a coding agent and it works immediately.
Matt Turck [42:05] The world's your oyster. There's so many things that you can just do. And you can be very creative and say, like, what do I always want to have, and it wasn't there? Nobody built this product. Can I build it now? And I think that mindset gives you useful things that make you more productive, but it also flexes your muscle. And sometimes it doesn't work. And then you understand, like, okay, agents struggle with this, or this is what I still need to learn to build these kinds of things.
Matt Turck [42:36] And I think that is the more intuitive perspective that's very useful, and that quickly gets you started on the path where you say, like, first there's excitement, then it's the sober reality, but then you pick up again and realize, okay, if I do it like this, I get more and more productive day by day. That's the more intuitive part. The more nuanced part is the part that I learned in the automation industry.
Matt Turck [43:06] I worked for, like, three years in the automation industry in Germany, automating factories, and it's a very calculated approach where you look at how you work, you time each of these steps, and then you say, if I automate this exact step in such a way, what could be the payoff? How much time would I save? And then you calculate what is the productivity gain, and then you calculate how much time do I need to develop this automation. And if you do that, you can quickly realize that automating certain things will not make a difference.
Matt Turck [43:34] The blog post I mentioned: emails don't really work. And there might be other things. A big thing is always calendar invites. Nobody likes to create invites for a meeting. But then if you think about it, you're also very particular about meetings. Some days you want more meetings, you have a meeting day, or you say, I can put this in before lunch, and agents know that. And if you specify that to an agent, you could also just create the calendar invites, just meeting invites, and it doesn't increase productivity by much.
Dan: managing agents like junior teammates (tools, guardrails, leverage)
Matt Turck [43:52] And so there are a lot of problems if you think about this nuanced way and can say, like, okay, I don't do that, but this will help me.
Tim Dettmers [43:56] And Dan, on your end, what have you learned or observed in terms of agents?
Matt Turck [43:57] What works?
Tim Dettmers [44:00] What currently doesn't work but will work soon?
Matt Turck [44:01] How to manage them?
Tim Dettmers [44:23] I think there's two broad things I've noticed for agents. The first is making the agents effective ends up being a lot like managing junior folks on your team or at a company. For example, the new intern who shows up on your team, you're not going to go to the intern and say, hey, go fix our revenue for the year, double our revenue for the year, or something like that. Maybe you'll try that once, but you're unlikely to see the payoff from that.
Tim Dettmers [44:48] Instead, what you often do with junior folks is you say, hey, here's a first little task that you can do to get to know this complicated codebase. Here are the things that you might run into because you've done it before. When you give the agents that context, give them that ability to look at those things, then they can usually figure things out. The other bit is that when you have a new person on your team, you maybe won't give them access to all the production credentials and all the production database and all those things, but you're going to give them enough tools to be productive.
Tim Dettmers [45:25] Sometimes there's this tension between, oh, I don't want my agent to go delete everything in production, so I'm just going to have it be hamstrung and watch every little thing it does. Whereas if you did that with a person, you would never expect that person to be productive. That's another bit. You want to think about the agents, at least today, as maybe interns or more junior folks. The other really interesting thing that I've noticed, when I think more about the educational role of a professor, is how do you prepare people for this future where agents are going to be such an important part of workflows?
Tim Dettmers [46:06] How do you train for that? One of the things that I've noticed is that the more expertise somebody has, whether that's in, for example, process automation or expertise in what I do in kernels and writing these very highly specialized programs, the more expertise you have, the more powerful the agent makes you. That's because you can work at such a higher level of abstraction. You know what the important things are. You know how to set the direction. You know what the common pitfalls are.
Tim Dettmers [46:25] What is easy? What's hard? What do you need to break up into multiple steps? One piece of conversation that was coming up for a while was, are agents going to replace all software engineers, or things like that? Or are they going to replace all junior people or something like that? I think where we are today, that's probably clearly not the case anymore. If I have a tool to make my team 10x stronger, I'm not going to fire nine people on my team.
Tim Dettmers [46:59] I'm going to say, okay, go do this, become 100x more productive than you were before. That's one bit. But then also the script for how you become an expert at something is probably pretty similar to the way that it was before. You're going to study things deeply. You're going to try to understand things a lot. You're going to want to do things yourself, get your hands on, and really get things done. In this world, ChatGPT can teach you a lot of things.
Tim Dettmers [47:24] Personally, I was trying to get ChatGPT to teach me all the little ins and outs of how a car works. I don't know how effective it was so far, but it's a lot easier to learn things now than it was certainly even two or three years ago. Those are some of the things I'd say. You want to treat the agents as if you are in that manager role. You want to help them get unstuck. You don't want to just say, throw the agent at the problem, walk away, and never look at it again.
Tim Dettmers [47:48] But you also want to figure out how to level yourself up so that you can be a better manager, have more domain expertise, and really understand things in a deeper way. The fact that you need to learn and be an expert, that doesn't change. I think that's very interesting, and that makes a lot of sense. The question is, if you show up on the job as a young kernel engineer and that's your first day, typically they would be, okay, well, you do this simple task and that other simple task, and then by year two, you graduate to a more complex task.
Tim Dettmers [48:05] How does hands-on job training look like?
Matt Turck [48:05] Right.
Education and training: learning in an agent world
Tim Dettmers [48:22] Yeah. So we think about this a lot together because we're still hiring aggressively, even in this world where the models are very good and the agents are very good. The way we think about it is that the professor in me went and actually recorded a bunch of lectures on how GPUs work. I make everybody watch those, and then I still give them a task from scratch, which is, okay, go take this FlashAttention kernel and modify it to do some other thing.
Tim Dettmers [48:58] You can pick the extra feature. The nice thing about the agents is that you can dive into that higher-impact role in a way that you weren't able to before. It's really impactful when a junior IC goes to try to manage someone for the first time because you're suddenly starting to think in much more precise terms. The classic software engineer thing is, hey, the PM asked me to do this and wrote this super long doc with all these requirements. But then the minute that you try to go ask someone else to do something, you realize the specificity that you need when you want to address a feature or something like that.
Tim Dettmers [49:28] The nice thing about agents is you can almost start to shortcut that process, where you can have the junior IC still be an IC, still do IC-style contributions, but they can now act in that manager role and act as their own PM. When you're communicating with the agents, you need to be as precise about what you're trying to get done. In some sense, I've seen with the junior folks who are joining my team—these are folks fresh out of college or fresh out of a master's—when they are really gung-ho about understanding and being able to use the AI agents, they're able to communicate so much better than in the olden days.
Tim Dettmers [49:56] They're able to level up their level of understanding a lot faster. And then, of course, they can do things and build tools at a speed that would have been really, really hard to do five, 10 years ago.
Matt Turck [50:23] And maybe I'll also add the educational perspective there because I think it's quite interesting. It's a little bit contrasting. What is also quite interesting is the educational perspective of using agents. I talk quite a bit about this: basically, use agents or be left behind. And that is also true for students. But just like as Dan said, you need the domain expertise. You already need to have some knowledge to use agents well. What we're seeing is, if we allow students to use agents, they are very productive, but sometimes they build solutions that look correct that are actually very bad or just wrong, and they don't realize.
Matt Turck [50:58] We are at this point where it's almost very difficult to learn both domain expertise and agent use. That's a very difficult balance to achieve because we don't want to have students that don't understand things, but we also want to have students that basically can use agents. And so, if they can't do that, they will not be effective in the workforce. What Dan said is, you already have a pretty good person with strong background knowledge, and then that person can level up their equipment with agents.
Matt Turck [51:26] But what do you do if you have someone that is just learning computer science? How much agents should they learn? How much work should you do without agents? And that's a very tricky balance, and we don't know how to solve it. If we let people use agents, they perform very poorly on basic knowledge. And if we let people just do the basic knowledge, they don't know how to use agents and they can't compete, so they can't be useful in the workforce nowadays.
Matt Turck [52:02] Maybe the solution is to do all sort of the basic knowledge first and then agents, but that's not what students do. Students have access to these AI tools. They will use them because it's easy. And so maybe the solution is just, you need a way of thinking and of working with information and knowledge that you don't understand. I guess, critical thinking. I think this goes beyond critical thinking. You basically need to know the unknown unknowns, things that you didn't consider and don't understand and that you didn't even think about.
Matt Turck [52:29] You need to have the ability to think more about that to really keep up with agents. Because I think in the future it's realistic that we work on problems that we don't understand, that agents understand, but we need to keep up in some way. That would be difficult. All right.
Tim Dettmers [52:36] So, to switch tacks as we get closer to the end of this conversation, what are you guys currently working on?
What Tim is building next (open-source coding agent; private repo specialization)
Matt Turck [53:04] What's top of mind for you at the Allen Institute on the one hand, Together AI on the other hand? Whoever wants to take it first. We have actually a very exciting project that will be released very soon, in the next weeks. And so I worked quite a bit on efficiency. I've been switching my work basically to coding agents. And so we will have a major release of an open-source coding agent that has a couple of key features. For one, training is 100 times cheaper.
Matt Turck [53:28] You need to generate synthetic data and you need to train on it. And so we have a method that's roughly 100 times cheaper, but still gets state-of-the-art performance. And then we have another major result, which you can almost see as the holy grail of open-source models. And that is, we can take a private codebase. Like, you have a company, you have this codebase, Claude doesn't know your codebase, but you have the data, you could fine-tune a model on it.
Matt Turck [54:00] So what we have is, you can just point our method to that repository. You don't need to have any sort of tests or need to understand how to generate the data. It's just automatic. You quickly generate the data, and then you have an agent that is as good as a frontier model, but you can have, like, a 32-billion-parameter agent. You can deploy it locally. You can have an army of specialized models for particular tasks, for particular codebases, and so forth.
Matt Turck [54:33] And yeah, I think that is a very powerful result. All of that is also packaged with a science of coding agents. There are a lot of confounding factors, a lot of hidden things in papers that are not mentioned. We sort of unearth them, build scaling laws, and show what does matter and what doesn't matter. And if you put all of that together, I think: the cheap agents, very few GPUs, everybody can use it. And very easily, we unblock people by revealing all the secrets.
What Dan is building next (inference efficiency, cost, performance)
Matt Turck [54:44] I think there will be a vast change in terms of how quickly we can progress in coding agents. Very cool.
Tim Dettmers [55:08] What about you, Dan? At Together AI, I think the major question that we're trying to answer today is: we have all these powerful AI models that can do all these amazing things. They're very expensive to run today. The question in all the public markets is, is OpenAI going to be able to...? And what's really exciting about Together AI is that we are kind of on the forefront of getting these models to use the hardware as well as it can. We talked a little bit about training early on in the podcast.
Tim Dettmers [55:29] At inference time, when you have the model, when it's already been trained, already been post-trained, the hardware utilization is less than 5%. So it's at a place where there's so much more that we can do. There's so much more that we can do to improve the performance efficiency that we can push out of these. We're really excited about figuring out how to use the hardware the best way that we can, whether it's serving customers like Cursor or whatever the next great foundation model company is.
Mega-kernels + Together Atlas (speculative decoding + adaptive speedups)
Tim Dettmers [56:07] We're really excited about pushing that frontier and then really getting us to this future that, if these models are going to be as impactful as you think, if they're going to be running everywhere, running your daily lives, running your toaster, well, we better make it the best possible toaster that we can. Do you want to talk about Mega-kernels and Together Atlas for a minute? Mega-kernels and Together Atlas are both projects along these lines. Let me dive into Mega-kernels first. To understand this, the first thing, when we say kernels, is we usually mean we are going to write a specialized GPU program for a single operation in a model.
Tim Dettmers [56:42] You can think of a model as one of these trained models as a bunch of different operations in a row, and there'll be hundreds of these. The way that we've been writing kernels for the whole history of NVIDIA hardware is that you really specialize a single kernel for a single operation. With these Mega-kernels, we're doing something quite interesting, which is we can take the entire model, however many billions of parameters, and put it into a single GPU kernel. With that, you can start to do a lot more fine-grained optimization than we were able to do before.
Tim Dettmers [57:12] It actually starts to make the NVIDIA GPU look a little bit more like a Cerebras chip or look a little bit more like a SambaNova chip in terms of the optimization that you're able to do. This is really critical at inference time. We're able to see 2x, sometimes 3x speedups over even highly optimized inference engines. We're working on bringing that to really work in production, bring it to fruition, and use it across our whole stack. Together Atlas is another really great, interesting research project that we've done recently where we can get the model.
Tim Dettmers [57:42] There's this technique called speculative decoding that we use at inference where basically we have a little model that is trying to guess what the big model is going to do. Because of the ways that we've designed language models, if the little model guesses correctly, you basically get those tokens for free. If you do the speculative decoding right, you can get 2x, 3x speedups over just running a vanilla model. With Together Atlas, we do one extra thing, which is we say, we can get that little model and we can actually adapt it to your traffic.
Tim Dettmers [58:07] The longer you use this model, the more it learns the patterns of what you're asking, of what you're saying, and it can actually get faster over time. So all these things, we have efficiency in mind, inference efficiency in particular. We're all pushing towards making all these things faster.
Predictions for 2026: small models, open-source, hardware, modalities
Matt Turck [58:31] What are you excited about for 2026 at a reasonably granular level? What do you think happens? What do you think doesn't happen? I think I'm sort of split. I think a lot of things will be very boring and there won't be much innovation, but then we're also surprised by a couple of things that we maybe don't see. And I think actually, sort of at the frontier, we will be less surprised. I mean, it's no secret that we ran out of pre-training data.
Matt Turck [58:57] And as Dan said, these are sort of the models, then you can sort of smooth over, and you smooth over with synthetic data. And that's how you build coding agents on lots and lots of different environments. You combine the data. We may make some progress there, but I think you already see the diminishing returns. I don't think coding agents will be that much better. The user experience will probably improve, but you see that all these models get almost equally good.
Matt Turck [59:25] GPT-5, and then later realize, oh wait, I used a different model because they're quite similar. And so I think we see less progress there. Where I think we see more progress is actually the small models. If you train smaller models on more specialized data, they can do quite well. And the smaller models that you get, they're pretty powerful. A 100-billion-parameter model, you can fit it pretty well, even on a sort of low-grade data center GPU, like an RTX 6000, which costs $6,000.
Matt Turck [1:00:03] I think for a lot of companies, it'll be very interesting. They don't need to rely on the frontier models. The small models might even be better because they're specialized. The big problem, and the Anthropic CEO pointed it out, is you have these powerful open-weight models, but nobody uses them because the deployment is so complex. And that is because once you go beyond eight GPUs, you first need the users to make it efficient, but then also very complicated inference systems. There's no open-source system that can do that at the moment.
Matt Turck [1:00:32] You need disaggregated inference, separation across sequence lengths, and so forth. Perhaps we can build this. We can build this also for an eight-GPU machine for smaller models. And then the efficiency that you see with a 100-billion-parameter model will rival what frontier models have. So you will get the efficiency of small models. You get the flexibility of small models. Performance on the frontier will stagnate, while on the smaller level, we get more and more powerful models still because you can distill from these large models into these small models.
Matt Turck [1:00:44] Taken together, I think that will change things.
Tim Dettmers [1:01:15] I'm also really excited about small models. I think we're going to see a lot more capability out of them. I'll be watching the open-source models pretty closely. GPT-7, you're starting to see it. You're starting to see the open-source models rival some of our current best frontier models. I think we're going to see another big jump in open-source capabilities this year. I'm really excited to see new hardware. We're starting to hear a little bit about Rubin, the next generation of NVIDIA GPUs.
Tim Dettmers [1:01:46] I think we're starting to hear a bit about the AMD 400 series of GPUs. I'm really excited to see what that next jump in hardware capabilities is, even as we haven't fully used even the current generation of hardware. I'm excited to see what people do with all the other modalities. I think last year, video generation models had a little bit of a moment with Sora 2, with Gemini and Veo, I think they called it. Really excited to see what they can do.
Beyond transformers: state-space and architecture diversity
Tim Dettmers [1:02:16] Really excited to just see what is that frontier of intelligence that you can get on your laptop or on your phone, and how fast can you push it? How far can you push it? I think it's never been a more exciting time to work in AI. You both mentioned state-space architectures earlier in the conversation. Do you think that's part of the near future, that we sort of evolve to post-transformer architectures with state-space, JEPA, world models, whatever direction? Is that something that you see in the near-term horizon?
Tim Dettmers [1:02:45] And that you think is desirable? I think in a lot of places, they're already there. Some of the best audio models in the world are at least partially based on state-space models. I think NVIDIA released a bunch of really great hybrid models recently. I think NeMoTron is what they called them. There's a lot of really great work there already. I think we will see the architectures continue to evolve. In some sense, the DeepSeek MLA compression takes some of those ideas.
Tim Dettmers [1:03:13] One of the MiniMax models had a linear attention idea. I think you're going to see a lot more diversity in architectures. You're kind of already seeing it, but certainly, I think, out of the Chinese labs, where there isn't really an OpenAI of China. There isn't an OpenAI or Anthropic or a Google Gemini that brings all these centers of product and model and revenue and all those together. I think you see a lot more risk-taking out of the Chinese labs, where you're trying to differentiate the next model, your next open-source model.
Wrap
Tim Dettmers [1:03:36] One way to do that is architecture. Another way to do that, of course, is just pure quality. I think we're going to see a lot more explosion of different architectures. All right. Well, this has been a fascinating conversation.
Matt Turck [1:03:40] I really appreciate the time and insight and thoughts.
Tim Dettmers [1:03:42] Thank you so much. Really appreciate it.
Matt Turck [1:03:43] Yeah, thank you so much, Tim.
Tim Dettmers [1:04:04] Thanks so much for having us. Great to see you, Tim. This was a lot of fun. Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already, or leaving a positive review or comment on whichever platform you're watching or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.