The Biggest Chip Ever Built — Why OpenAI Runs On It | Cerebras CEO Andrew Feldman
The MAD Podcast with Matt Turck · with Andrew Feldman, Co-founder and CEO, Cerebras
Andrew Feldman is the Co-founder and CEO at Cerebras. We cover why tokens per second per user is the key inference metric, how Cerebras’ wafer-scale chip is 58 times larger than a GPU with thousands of times more memory bandwidth, and why generating one word can require moving model weights equal to 100 HD movies.
Chapters
- 1:31 — Why speed became the AI bottleneck
- 2:32 — Tokens per second per user, explained
- 3:16 — AI’s broadband moment and the Netflix analogy
- 4:35 — The AI chip landscape: GPUs, TPUs, Trainium, ASICs
- 6:36 — What is an ASIC?
- 8:08 — Nvidia, Groq, and the fast inference war
- 9:16 — OpenAI, Broadcom, and specialized silicon
- 12:10 — China, power, and sovereign AI infrastructure
- 15:05 — Is the AI infrastructure boom a bubble?
- 18:56 — The hidden bottlenecks: HBM, CoWoS, and 3nm
- 22:57 — Why agents are creating CPU demand
- 25:36 — Andrew Feldman’s path from SeaMicro to Cerebras
- 26:13 — Why Cerebras bet on AI in 2016
- 31:14 — SRAM vs. HBM: why inference is a memory problem
- 33:19 — What wafer-scale computing actually means
- 34:28 — The deep-tech “Everest” problem
- 36:07 — The moment the first Cerebras system worked
- 36:49 — Ringing the bell and surviving deep tech
- 39:08 — How a giant chip handles failure
- 41:22 — Why GPUs struggle with decode
- 42:17 — Prefill vs. decode explained
- 44:01 — The “100 HD movies” problem in AI inference
- 45:04 — How fast inference changes RL and training
- 48:08 — Reasoning models and why they cost more compute
- 50:08 — Verification, guardrails, and small models checking big models
- 52:37 — Multimodal AI and the path to video
- 53:51 — Cerebras’ business model: hardware, cloud, and API
- 55:14 — OpenAI’s 750MW inference deal
- 55:36 — Why data centers are measured in megawatts
- 58:01 — AWS Trainium + Cerebras decode
- 59:29 — Fast tokens as a cloud product
- 1:00:52 — Is CUDA still a moat?
- 1:03:53 — How TSMC helped Cerebras build the giant chip
- 1:07:41 — Why nobody cared in 2020
- 1:08:15 — Why chip supply chains are hard to diversify
- 1:09:54 — Why today’s AI models will be the worst you ever use
- 1:10:38 — What fast AI could do to SaaS
Transcript
Why speed became the AI bottleneck
Matt Turck [1:37] I thought a fun place to start would be to talk about speed. So, has speed become the dominant conversation for AI today?
Andrew Feldman [2:06] Yeah, what happened, I think, was for a long time AI was sort of a novelty, right? It was like a parlor trick. It was cool, but not useful. And what happened somewhere around the middle of 2025 was the AI got smart enough such that people began to use it. And remember, we make AI with training, but we use it with inference. And suddenly people wanted to use it. And the minute you want to use it, the minute it's productive, right?
Tokens per second per user, explained
Andrew Feldman [2:32] Speed matters, right? Fast tokens are more productive. And so the conversation moved from everything else to how do we make our inference faster? How do we deliver tokens more quickly? Because those are more productive tokens. We get more done in less time, and therefore it's more valuable.
Matt Turck [2:40] Right. And what does speed mean? Is that a question of token speed? Is that completion of the task? What's the right metric?
Andrew Feldman [3:15] The right metric is tokens per second per user. That's how fast you get the first token all the way through the last token in your response. And it's true for queries from chat, but it's also true for agentic flows, right? If they're sort of multi-cycle turns, waiting is amplified. And so what you want is blisteringly fast responses so that the AI feels like it's in real time. You can engage with it.
AI’s broadband moment and the Netflix analogy
Matt Turck [3:20] So it's the broadband moment for AI.
Andrew Feldman [3:50] I think that's right. And I think that's a very good analogy. I think if you think of something like Netflix, right? When the internet was slow, right, Netflix delivered DVDs in envelopes. You would get a DVD in an envelope. And when the internet became fast, they didn't get more efficient at delivering DVDs in envelopes. They became a movie studio, right? The speed enabled them to become something completely different. And that's what speed does in general and in particular for AI.
Andrew Feldman [4:06] It opens up a whole new domain. It allows you to use the AI differently. You will stay longer, you will come more often, and you'll work on harder problems.
Matt Turck [4:12] Yep. So it's literally a question of the UX, right? That's just—nobody wants to wait a few seconds.
Andrew Feldman [4:26] Yeah, that's right. I mean, how big is the market for slow search? How big is the market for dial-up? It's zero. How long will you wait for a website to resolve? Will you wait eight seconds? Nobody waits. And so it's the exact same with AI.
Matt Turck [4:31] Yep. So no more people waiting with their laptops open while the agent—
The AI chip landscape: GPUs, TPUs, Trainium, ASICs
Andrew Feldman [4:36] That's right. While it's running and running and running, I think that is not what people want.
Matt Turck [5:16] Okay, great. Wonderful. So I'd love to talk about the landscape of the chip industry right now to help people visualize, in part, where you guys are. So there used to be basically this concept of one chip to do it all. And obviously this is evolving dramatically. There's you guys, there's Groq, there are GPUs, and people have heard of Trainium, people have heard of TPUs. So help us compare and contrast who does what. What for what?
Andrew Feldman [5:54] Well, I think there used to be one chip to do it all. It was called a CPU. Yeah, right. And there emerged a coprocessor to do discrete graphics. And as the AI workload became interesting, we focused chips more and more on that particular workload. And today, several companies make traditional GPUs, so NVIDIA and AMD. They make very standard GPUs. There are a group of companies—the hyperscalers—that make some of their own parts. So, the TPU.
Matt Turck [5:56] So TPU is Google.
What is an ASIC?
Andrew Feldman [6:36] TPU is Google, and Trainium is AWS. And then there were a group like Cerebras, and we were among the pioneers to build a part from scratch optimized for AI and nothing else. And that we weren't inside of a hyperscaler, we weren't optimized for a hyperscaler's problem or for one lab's problem. We were building a chip for a collection of AI problems, and all our thinking was around how to accelerate AI. And that's sort of the landscape to this day.
Matt Turck [6:50] Yeah. There's even more specialized ones, right? Like Etched, developed specifically for transformers. Those are called ASICs as well. Could you maybe define what that term means?
Andrew Feldman [7:30] An ASIC is an application-specific integrated circuit. And it's a word that now has a wide range of meaning. It means you have made a series of choices away from general towards a narrower class of problem solving, and that you've made some decisions in your architecture that make it much better at some things and much, much worse at others. Right? And that's choices that are made across the spectrum. So the TPU has made some choices like that. It can't do graphics.
Nvidia, Groq, and the fast inference war
Andrew Feldman [8:09] It's very good at matrix multiply. It's not very good at a collection of other things that we do in mathematics. Same for the GPU. We've all made different choices. Right now in production, there are NVIDIA and AMD. There are a collection—there's the TPU for Google, Trainium. Just coming up is a Maia part, is a part from Microsoft, Cerebras, and there was one other: Groq, that got acquired by NVIDIA. Yeah.
Matt Turck [8:23] And to the general chip versus specialized chip, it was actually interesting because you guys and your CTO celebrated when Groq was acquired. What was that? Was it a recognition by NVIDIA that your vision was right all along?
Andrew Feldman [8:59] Yeah, I think one of NVIDIA's sort of most durable moats was the perception that the GPU could do everything and it was all you needed for AI. And the acquisition of Groq for $20 billion and the structure and the speed with which they chose to do it made clear to everyone that that wasn't true, that the GPU architecture could not do fast inference, and that this market was large and growing quickly. And we were the fastest at it and the largest.
OpenAI, Broadcom, and specialized silicon
Andrew Feldman [9:17] And our sales were more than 10 times Groq's, and they paid $20 billion for the number two player. So that was a good day. That was a good day.
Matt Turck [9:33] And the recent announcement of Jalapeño between OpenAI, which is a very large customer of yours, and Broadcom. Where does that fit in that picture? Is that one of those highly specialized ASIC types?
Andrew Feldman [9:33] Right. Groq.
Matt Turck [9:34] And it's for inference.
Andrew Feldman [10:05] That's right. Jalapeño is a part that has been a long time coming. It was announced with Broadcom before OpenAI did the deal with us. Remember, we did a huge deal with OpenAI. This is probably one of the largest deals in Silicon Valley history, north of $20 billion. I think they have a yawning need for silicon. And one of the things that OpenAI has been, I think, the best in the industry at is looking at an exponential curve adoption rate of growth and not being afraid of what it says.
Matt Turck [10:24] Right.
Andrew Feldman [10:54] Others have had to go out and strike really bad deals to get capacity, where OpenAI saw this coming. They struck big deals for memory, for compute with us, with others. They've really been sort of visionary in understanding what it means to extrapolate from an exponential curve. And that curve is AI usage, right? It is growing so unbelievably fast. The whole industry is chasing demand.
Matt Turck [10:55] Right.
Andrew Feldman [11:09] We usually—well, often it's the other way. Often people are building it, hoping it will come. And in our case, all of us are chasing what people already want to do, let alone what they might do in the future. Hey, Ed.
Matt Turck [11:30] As a point of opinion, Andrew, I mean, because that's such an interesting topic, to finish on GPUs versus inference versus ASICs, is that the permanent situation from your perspective, that we're going to be in this forever multi-silicon kind of environment?
Andrew Feldman [11:55] Yeah. Multi-silicon environments are a healthy ecosystem, right? I don't think anybody would say that the x86 environment was a healthy place. For 20 years, there were two players. And if you notice, when a new workload came around, the cell phone workload, that was very similar but required low power and battery, both lost, right? Intel got zero share. AMD got zero share. And a company nobody had previously heard of became the largest seller of compute in the world, and that's ARM.
China, power, and sovereign AI infrastructure
Andrew Feldman [12:10] And that healthy ecosystems have lots of different ways to solve problems.
Matt Turck [12:29] Where does China fall in all of this? So there's the Huawei Ascend chips for DeepSeek and this emergence of a full-stack Chinese AI factory, for lack of a better term. Is that divide happening, or is that overstated?
Andrew Feldman [13:01] I don't know that that divide is real. I think they are an industrial adversary. I think they have made really interesting investments. They've made investments in power, in their grid, which is a real weakness in the U.S. I'm sure in France you have nuclear, which sort of turns out to be pretty cheap power and pretty clean. But in China, they've made huge investments in their grid. They are behind in chips, but their approach at the next level is open-source models, where they're producing some extraordinary models.
Andrew Feldman [13:18] Not as good as GPT or Anthropic or Google's Gemini, but very good.
Matt Turck [13:18] Right.
Andrew Feldman [13:32] And I think they don't have the same chip strengths that the U.S. does, but they've got leadership in other domains, particularly in power, which is what we need for data centers.
Matt Turck [13:37] Okay. Where does China fall in Cerebras's strategy?
Andrew Feldman [13:38] We don't sell.
Matt Turck [13:44] For geopolitical reasons, for business reasons, for regulatory reasons?
Andrew Feldman [13:51] For regulatory reasons and for some geopolitical reasons. Okay.
Matt Turck [14:09] Is there a role for local AI and local chips? NVIDIA had some announcement around just building chips for Windows computers. Is that for inference? Is that something that you guys think should be part of the multi-silicon ecosystem?
Andrew Feldman [14:35] Yeah, I think if you look at the way the ecosystem for apps and cloud emerged, everything you can do, you should do on your phone or your laptop. But the ability to get real processing power to a phone or to a laptop is constrained because they're generally working off a battery. And so you want to do as much as you can as close to the data as you can. But the truth is, in most situations, for real compute, you have to go up to the cloud, to the data center.
Andrew Feldman [15:01] And that's exactly the way it's going to be with AI. We're going to do a lot of work on the cell phones, on the laptop, but for the big work, you're going to go to the data center. And that's where our focus is. Our focus is data center compute for AI.
Is the AI infrastructure boom a bubble?
Matt Turck [15:28] Great. So you mentioned how powerful the demand was, and demand ahead of production. So, just to ask the inevitable question around market equilibrium: bubble or not? Five trillion in market value was lost, and NVIDIA, I think, lost $300 billion that day. Now, that was corrected within 48 hours. So the market seems very jittery around that general concept. Maybe, without putting words in your mouth, just based on what you said, though, it sounds like you're reasonably on the other side of that concern.
Andrew Feldman [16:24] Okay. I think that if you spend a lot of time staring at the public markets and watching their ups and downs every day, you're watching the wrong thing. The wisdom received from the great thinkers of public market investing is that, in the short term, the public markets are voting mechanisms: who's most popular? But in the long term, they're weighing mechanisms: who's created the most value, the most weight? And so what we can do as an early public company is, we cannot look at the day-to-day fluctuations, and we focus on just building extraordinary technology, winning new customers, making our existing customers happy, ensuring they buy more, being better every week.
Andrew Feldman [16:53] And so I think the day-to-day fluctuations are out of our hands, and that's about the voting.
Matt Turck [17:25] Yeah. And maybe just to play devil's advocate on the demand, it seems like a lot of the demand comes from the big labs, which themselves are financed by venture capital, private equity, hedge funds, whatever you call it, sovereign investors. Is there any concern that demand for chips comes from the labs, which are maybe artificially financed, and that if you had a few Circle deals here and there, don't you sense any fragility?
Andrew Feldman [18:02] That's not exactly our experience. I mean, obviously we have enormous demand from OpenAI, but we have huge, dozens of other customers who are trying to place very big orders on us. And historically, bubbles were when supply got out ahead of demand, right? In the '90s, we built out telco infrastructure, right? We built out fiber years before it was going to be used, and it took six or eight years, and it all got used. But it was a sort of, if you build it, they will come mentality.
Andrew Feldman [18:44] Whereas what's different about AI right now is we're all trying to catch up. We're trying to build data centers faster. We're trying to increase our supply chains for demand that's already here today. And only a very small portion of the world are using AI anywhere close to its potential. And we're already sort of overwhelmed with compute. We're overwhelmed with the demand for memory, which is a real weakness in the GPUs. It's not a problem we face.
Andrew Feldman [18:54] And the ecosystem, there are constraints left and right. And that doesn't feel like a bubble.
The hidden bottlenecks: HBM, CoWoS, and 3nm
Matt Turck [19:07] Do you want to actually double-click on that? Because that's so interesting, right? It seems like the bottleneck keeps moving. So what's the current problem of the memory shortage that people may have heard of?
Andrew Feldman [19:47] So there are three major bottlenecks right now. The first is all GPUs and all ASICs except us use a type of memory called DRAM and a particular flavor of that memory called HBM. And that memory is made by three companies in the world: SK Hynix, Samsung, and Micron. Three companies. And they're sold out. That's the number one problem. We don't use it.
Matt Turck [19:53] And it's sold out as in, like all things hardware, the only way to increase supply would be to build new factories.
Andrew Feldman [20:26] Build more factories. And so you have this huge problem for GPUs and even now for CPUs, in which you can't get this memory because of our architectural choices. We don't use it. The second constraint was a process at TSMC. The process was called CoWoS, and this used silicon as a motherboard. And this is what NVIDIA and AMD and others use as a motherboard. So they put their chips and the memory chips on it, and it's more efficient than a traditional motherboard.
Matt Turck [20:36] Go ahead.
Andrew Feldman [21:03] That process is sold out, is highly constrained. We don't use it. Third limitation in the space is 3-nanometer fab space at TSMC. So this is the most aggressive factory, with the most advanced technology. Our chips are at 5 nanometers. So we don't use the 3-nanometer technology.
Matt Turck [21:14] And that's so fascinating. So, to ask the layman's question, what does that mean, 3-nanometer factory? That's a factory that's solely dedicated to doing—
Andrew Feldman [21:44] That's right. They build factories, and the factory etches transistors that are a certain distance apart. And the smaller the distance apart, the more transistors per square millimeter, and the more you can do per square millimeter. And so the history of the computer industry making chips has been, we've gotten better and better at putting transistors closer and closer.
Matt Turck [21:45] Yeah.
Andrew Feldman [21:54] And so right now, the bulk of the GPU world is all at 3 nanometers, and there's tremendous congestion at the factory.
Matt Turck [22:10] And that's literally different tooling at the factory?
Andrew Feldman [22:18] Tooling at the factory. And we use the 5-nanometer factory, and this is our chip here.
Matt Turck [22:20] Yes.
Andrew Feldman [22:34] This is the largest chip built in the history of the computer industry. It's 58 times larger than a GPU. It has 2,500 or 3,000 times more memory bandwidth.
Matt Turck [22:36] Boom.
Why agents are creating CPU demand
Andrew Feldman [22:57] And for AI, bigger chips process information more quickly, and therefore you get answers in less time. And obviously, you don't want a bigger chip for a laptop or a cell phone or for lots of other things, but for AI work, big chips are undoubtedly the best way to go.
Matt Turck [23:23] Okay. Fascinating. All right, so I want to do a bit of a deep dive on that, on the chip itself, in a second. But to close on this, so, three bottlenecks. You also mentioned CPUs a couple of times in this conversation, and there seems to be a theme around the emergence of, like, a CPU shortage as well. Is that, one, true? Two, what causes it?
Andrew Feldman [24:00] So agentic AI is a world in which AI doesn't just provide answers, it initiates action. So that action might be: go to a website, learn, gather some data from a website, bring it back, take another action. Those actions are done by CPUs. And so, as AI gets better and better at doing things, at making instructions, calls for things to get done, we're using more and more CPUs, right? And that is driving up the consumption of CPUs and therefore the demand for CPUs.
Matt Turck [24:09] Okay.
Andrew Feldman [24:39] And so this sort of huge push for more CPUs is being driven by AI on machines like ours and GPUs doing agentic work and asking the CPUs to take an action: to go to a website, to order a burrito, to find a piece of information, to pull it from storage. All that work is being done by the CPUs.
Matt Turck [24:41] Including in a system like yours?
Andrew Feldman [24:43] Including in a system like ours.
Matt Turck [24:49] Oh, that's interesting. So what does it mean? So you'll get, you have your chip and CPUs on the side.
Andrew Feldman [24:50] Yeah.
Matt Turck [24:51] And those would be a system.
Andrew Feldman [25:15] It means that AI, when we do AI in agentic, the AI processor is like the brains and the CPUs are like the body. They're taking action, they're doing things in the digital world under the direction of the AI, which is running on the accelerator, on the Cerebras system. And so, as we do more and more AI work and more and more agentic work, we're making more and more calls to CPUs, and therefore the demand for the CPU is just through the roof.
Matt Turck [25:33] And CPUs experience the same memory shortage.
Andrew Feldman’s path from SeaMicro to Cerebras
Andrew Feldman [25:36] Same memory shortage. It's the exact same memory. Oh, that's right.
Matt Turck [25:55] Okay, yeah. Let's go back to the product in a second, but let's talk about you and your journey a little bit for people to have context. So, obviously, you pulled off an incredible IPO with perfect timing, and I know it was a little bit—
Andrew Feldman [26:03] The way you have perfect timing is to have horrible timing for 10 years, so you have perfect timing.
Why Cerebras bet on AI in 2016
Matt Turck [26:13] And very much to this point, so you guys started building this bigger chip, focused on inference, over a decade ago.
Andrew Feldman [26:14] 2016.
Matt Turck [26:28] 2016, which makes perfect sense in retrospect, given how long it takes to build a technology like this. But what was the vision? 2016, I guess it was like four years after ImageNet. Deep learning was a thing.
Andrew Feldman [26:30] It was all vision.
Matt Turck [26:31] It was all vision.
Andrew Feldman [26:51] It was all vision. And that's why we don't believe the right thing to do is to embed the latest and coolest model into your circuitry. That's a mistake. We're the fastest in the world at transformers, and our architecture was set before transformers existed.
Matt Turck [26:52] Right.
Andrew Feldman [27:07] We're the fastest in the world at diffusion, and our architecture was set before diffusion. What you want to do is get the underpinnings of those so that, when the market moves, you can be good at that as well. Otherwise, because of the delay, short life.
Matt Turck [27:13] To the ASICs question earlier, so that's what others do: they build the architecture of the model into the—
Andrew Feldman [27:17] Some have, some have. And historically, that's been a structural mistake.
Matt Turck [27:23] It's been a structural mistake. So you're doing the broad neural reasoning platform on which any kind of model would work.
Andrew Feldman [27:41] You want to think about what is the underlying calculation. The underlying calculation of all this work is sparse linear algebra. And if you can accelerate that, whatever the model builders invent, you can make faster. And that was our approach.
Matt Turck [27:52] Yeah. So you started in 2016. You had a prior company that you sold to AMD. What were some of the lessons you learned there that you took into Cerebras?
Andrew Feldman [28:22] I think the lessons are large and many. I think experience is another name for having made mistakes and learning from them, right? I think we, as a team, have built dozens of chips over the past 25 years, and the returns to experience in chipmaking are enormous. And we built a different type of computer at SeaMicro, a type of computer optimized for low power and optimized for a workload that was very different than AI, for something like web browsing.
Andrew Feldman [29:06] But the fundamental underpinnings, the questions you ask as a computer architect, are always the same. What can I do to make this work faster? And is there enough of it to make it worthwhile? These are the two questions we ask. Should we build a part for it? What could we do to build a chip optimized for AI? And will there be enough AI so that you can build a business around it? Those are the questions we asked in 2016.
Andrew Feldman [29:26] And the flip side of that was: wouldn't it be a surprise if the GPU, which had been optimized for graphics for 20 years, had been pushing pixels to a monitor, was suddenly good at a new world?
Matt Turck [29:28] It'll—
Andrew Feldman [29:46] Wouldn't that be serendipitous? And we came to believe that it wasn't the right architecture for it. It was just better than the CPU, and that we could build an architecture that would be vastly faster, that would use less power, and could drive down the cost of it.
Matt Turck [29:55] And that was the journey. So Cerebras was very early, but then we spent time in the desert. Okay.
Andrew Feldman [29:57] Then we wandered. Then we wandered in the desert.
Matt Turck [29:58] Yeah.
Andrew Feldman [29:59] Yeah.
Matt Turck [30:17] Maybe walk us a little bit through those years for the founders, especially deep tech founders listening to this. So first of all, what was the issue? Was it market timing? Was it that the technology was not working? And then how did you go about it as a team? I guess your board and your investors, and raising more rounds, as you presumably didn't have the proof points that you needed.
Andrew Feldman [31:01] So first, we, at the beginning, were honest with our VCs and told them we were going to attack a really hard problem. We weren't going to build something that was a little bit better than a GPU. And our idea, our strategy, was that you will never beat a great company like NVIDIA by doing something a little bit better than they do. They're going to buy everything for less. They're going to have pricing pressure. They're going to be able to bundle.
SRAM vs. HBM: why inference is a memory problem
Andrew Feldman [31:40] And the right strategy would be to do something incredibly hard in engineering that was way better—10, 15, 20, 30, 50 times better. But to do that, the ordinary and obvious paths are all closed. Everybody else has taken them already. And so what we observed was that speed in inference was going to be a function of memory, and that there were two types of memory. There's this DRAM, or HBM, and they can store a lot, but they're slow. There's another type of memory called SRAM that is unbelievably fast, but per unit area can't store very much.
Andrew Feldman [32:18] And that for graphics, everybody had always used DRAM. They used HBM, and the AI workload was fundamentally different. In graphics, you move data to the GPU and then you work on it for a long time, and then you send the result. So the time spent in total of movement plus work is dominated by work. In inference in AI, it's the exact opposite. You move a huge amount of data, all the weights from memory to compute, and you do one calculation to generate the next word.
Matt Turck [32:29] Mm-hmm.
Andrew Feldman [32:30] And then you have to do it again.
Matt Turck [32:31] Yep.
Andrew Feldman [32:40] So all the time is dominated by the movement of data. So that's why GPUs have so much trouble being fast.
Matt Turck [32:40] Mm-hmm.
Andrew Feldman [32:59] So we observed that if we chose a strategy using SRAM, we could be faster, but then we have to overcome the trade-off with SRAM that it can't store very much. That led us to the solution that if we could build a part vastly larger than any part in history—the size of a dinner plate—we could stuff it to the gills with SRAM and thereby get the benefit of SRAM, that it was fast, and overcome the weakness that it can't store very much.
What wafer-scale computing actually means
Andrew Feldman [33:19] And that led us to a strategy called wafer-scale. This chip is made from a single wafer. It comes out of TSMC.
Matt Turck [33:21] Do you want to define what a wafer is?
Andrew Feldman [33:56] A wafer—all chips are cut from a wafer. A wafer is a circular piece of silicon that's 300 millimeters across, and the process of chipmaking stamps out, like your mother does with a cookie cutter, stamps out chips. The biggest chip that had ever been built before us was 800 square millimeters, 840 to be exact, and this is 46,000. So we had to invent all sorts of new technology to build a chip this big. And once you build it, you have to invent ways to power it, to cool it. There are no vendors waiting for you.
Matt Turck [34:07] Yeah.
The deep-tech “Everest” problem
Andrew Feldman [34:29] Right. Because you look like nothing else ever made. And that took years. And it was a deep-tech problem. It had never been solved before. And we had an 18-month period where we were spending $8 million a month and we couldn't build them. So if you're deep-tech founders, you have $8 million a month.
Matt Turck [34:34] Yeah. Now, why so much? What was the cost? Just to understand how these businesses—
Andrew Feldman [35:03] Because what everybody thought was hard, we solved quickly. And what nobody else knew about, because they'd never actually done it, turned out to be really hard. Imagine, I tell people, imagine the first group that was going to climb Everest, and they're at base camp and they're having tea with a group that just failed. And the group that just failed says, halfway up, there's this part. It's unbelievably hard. We couldn't do it. Okay. Your team climbs up, makes it all the way to the top, comes back.
Andrew Feldman [35:24] They're having tea again. And the team that made it leans over to the team that hadn't made it and says, that part in the middle, that wasn't the hard part, because nobody had gotten past it. Nobody had gotten past certain things, so they didn't even know what to be afraid of.
Matt Turck [35:25] Success.
Andrew Feldman [35:46] We now know, and it was something—a step called packaging—and that's how you affix a wafer to a motherboard, how you deliver power to it, and how you cool it. And nobody had done it before. And over that 18-month period, we became the best in the world at it from approximately zero.
Matt Turck [35:46] Right.
The moment the first Cerebras system worked
Andrew Feldman [36:15] And we did that by failing again and again and using good engineering methodology and doing a failure analysis of every single failure. So we failed differently again and again and again and again. And we told our board—we met with our board every six weeks—this is the strategy. Nobody's ever done this before. This is what we're going to do. And we had to invent new materials. We had to invent new techniques. We ended up building things that everybody else had partners who could do.
Andrew Feldman [36:42] But it all gets to 2019, when Alex Wood solved it. And my co-founders and I, the first time it worked, were sitting in a tiny little lab, and we couldn't believe it. We were the first people in history to make one work. And we just stared at it. Watching a server run is about as exciting as watching paint dry.
Matt Turck [36:42] Yes.
Ringing the bell and surviving deep tech
Andrew Feldman [36:50] And there we were, the five of us just staring at this machine, not believing it might work. We'd made it work.
Matt Turck [36:56] Was that a bigger moment than actually ringing the bell, or just completely different things emotionally?
Andrew Feldman [37:13] Emotionally, it was a completely different thing. It was that we had made our idea work. And I think for deep tech founders in particular, there's always this little thing in the back of your mind that says, maybe it's shit.
Matt Turck [37:14] Right. Maybe we're actually crazy.
Andrew Feldman [37:43] That's right. Maybe it's going to fail, and maybe it's not going to work. And maybe I don't have time. Maybe we're going to run out of money. Maybe, maybe, maybe. And the flip side of that is the joy that this was my co-founders' ideas. I mean, these are their ideas manifest in the world, and that is an extraordinary thing. That is when someone's ideas take physical shape. And then the next step is when you watch other people's work sit on top of your idea. Then you know you love making infrastructure.
Andrew Feldman [38:25] And so when we rang the bell and we went public on May 14th this year in the largest semiconductor IPO in history, we did something unusual. We invited all our engineers who'd been with us since the start, everybody who'd been with us more than nine years, and their families, and we all rang the bell together. We all got up on stage, and that was sort of a moment of a different pride that we had done this together and that we had managed not to die.
Andrew Feldman [39:01] Sometimes with a startup, people aren't honest. They don't tell you that, but we'd avoided some. We'd made plenty of mistakes, but we'd avoided the fatal ones, and we'd made it through to a level of success that gave us the opportunity to pursue a new level of success. That's what an IPO is. It's not the end. You've gotten to a plateau that you can chase a new level of success in the public market. And that felt pretty good, I'll admit.
How a giant chip handles failure
Matt Turck [39:27] Amazing. Amazing. Thanks for sharing. So, just to go a little deeper on some of the product and technical stuff, what are the trade-offs of building a bigger chip? One that may come to mind is failure mode.
Andrew Feldman [39:27] Sure.
Matt Turck [39:38] If you have a lot of little chips on a big wafer, you can isolate the problems. If you have a big one, then everything will fail at the same time.
Andrew Feldman [40:06] So you have to think very carefully about failure mode, and we invented a technique that has about a million identical tiles, and if one fails, we can shut it down and use a redundant one. And they keep going. So if you're going to go big, you have to think about, in the very architecture of the computer, how you're going to manage failure. I mean, the GPUs have a huge failure rate, so I'm sure you guys have spoken about this.
Andrew Feldman [40:44] Infant mortality is enormous, and they fail all the time. There's good data. Facebook put out a paper on the number of failures they get in a big cluster. Now, we can do some other things because we have all this compute in one spot. We can invest more to cool it. So we pioneered water cooling in AI systems, and we run these much colder than GPUs. And the failure mode in electronics is temperature. And so, by running them colder, we are more reliable.
Why GPUs struggle with decode
Andrew Feldman [41:22] And so that was an advantage. But again, we had to invent the technique to cool off a chip this big. And so I think we have a systems mentality, right? If you're going to do something big, you've made trade-offs, and you have to think about an architecture that allows for redundancy and repair. You have to think about the pros and cons of every aspect of the architecture. And that's how you do something different and new. Great.
Matt Turck [41:45] And again, just to drive the point home, to make sure there's, I guess, a clear takeaway for people from this conversation: explain it to me like I'm maybe not five, but 15. Why is this faster than a GPU? Like, what fundamentally makes it faster, and what compute is super fast with a GPU?
Prefill vs. decode explained
Andrew Feldman [42:17] How tokens are generated in inference is why it's faster. In inference, to generate a word—and our answers are a whole stream of words; they can be code, they can be pixels, but call them words—the weights of the model are moved from memory to compute, a calculation occurs, and that generates the word.
Matt Turck [42:19] Is it prefill versus decode?
Andrew Feldman [42:21] That is both steps.
Matt Turck [42:24] Okay. And do you want to maybe explain what prefill and decode are?
Andrew Feldman [43:01] Okay. There are two steps in the computation necessary to do inference. When you type into ChatGPT, "Explain to me the history of this village prior to World War II," and it can't see you, two things have happened. The first thing is your prompt has been processed. That's step one. And step two: your answer has been generated.
Matt Turck [43:02] Sure.
Andrew Feldman [43:07] And the way your answer is generated is called decode.
Matt Turck [43:09] Okay.
Andrew Feldman [43:43] And decode is sequential. Processing your prompt, which we'd call prefill, can be parallelized. So you can process many of them simultaneously, but the speed with which you get an answer is a function of the decode, and it is step by step by step in sequence, and that can't be changed. And how you do that step is you move weights, which are the intelligence from the model, to compute to generate a word. You do a calculation and you get a word, and then that word is used to generate the next word, where weights are moved from memory to compute.
The “100 HD movies” problem in AI inference
Andrew Feldman [44:22] So the process is one of moving weights from memory to compute. How big are the weights? In a little model, like a 70 billion parameter model, the weights are about the size of 100 HD movies. So to generate a single word, you move 100 HD movies from memory to compute, and then you have to move them again for the next word.
Matt Turck [44:23] Hmm.
Andrew Feldman [44:52] And you want to do this 1,000 times to get a good answer, a 1,000-word answer. This is where HBM is slow. This exact step is where HBM is slow. And that exact step is where, by having all the SRAM here, we're blisteringly fast. And so the speed of moving weights to compute is about 2,500 times faster here than on an NVIDIA GPU. And so that's the essence of what we've been able to do here and why it's so much faster.
How fast inference changes RL and training
Matt Turck [45:14] Fascinating. Do you think that models need or will evolve as well in that fast AI world, or is that just a chip problem?
Andrew Feldman [45:35] No, exactly. Because remember, two things are happening. One, we want the models to be smarter. Two, one of the ways models are getting smarter is with RL, and RL uses inference inside of training. And so the faster you can do the inference inside of training, the faster you're training.
Matt Turck [45:37] Interesting. Is that a market for you guys in AI?
Andrew Feldman [45:39] Yes, that's a market for us as well.
Matt Turck [45:42] So you're not just inference, you're also—
Andrew Feldman [45:52] We do RL and we do traditional training too. Not for the largest models, for the largest labs, but for the next tier.
Matt Turck [45:56] For pre-training, for pure pre-training?
Andrew Feldman [46:00] Pre-training, fine-tuning, our full set.
Matt Turck [46:12] So you can do pre-training, I mean, post-training with RL, some pre-training for the other labs. GPUs are still better than you for what job?
Andrew Feldman [46:45] GPUs in training have some challenges that have been solved by a very narrow selection of the community. A GPU is a very small chip, and the calculations that we need to do in training are very large. And one of the most complicated parts of training is breaking up the calculations and spreading them apart on multiple GPUs. And that's called distributed compute. That has historically been the domain of the supercomputer world and is very difficult, not just because cracking a problem and having lots of others work on it is hard, but they have to constantly share information.
Andrew Feldman [47:38] And that sharing information is why NVIDIA needed to buy Mellanox, right? They needed to control a fabric over which all this sharing, in order to get an answer, would happen. Breaking up a big matrix multiply, a big calculation, is called running tensor model parallel. You are breaking up the tensor and spreading it apart. And the best labs in the world are good at that, but nobody else is.
Matt Turck [47:39] Right.
Andrew Feldman [48:07] When we run training, we don't have to run that way. We run what's called model data parallel. And data parallel is very simple. And so it allows teams who are good, and very good, to quickly test ideas in training. And so we are easier to use and faster because we allow them to use a technique which is much simpler.
Reasoning models and why they cost more compute
Matt Turck [48:30] Let's talk about the agentic world, maybe starting with reasoning. I think I saw a blog post where you guys, again, stated that reasoning was not always the right solution for all problems. How do you think about this, and where does that fit in the agentic world?
Andrew Feldman [48:54] We're using it as a technique, a little bit like when you were in eighth grade and you wrote different drafts of the paper. Single-shot inferencing: you write a query, you get an answer. The easiest way to think about reasoning is you're going to do several graphs. It's going to break the problem up, it's going to solve it in parts, it's going to bring them together, it's going to review the results, it's going to improve the results, and then it's going to give you an answer.
Andrew Feldman [49:44] And so that is going to take more compute. And if your compute is slow, that's going to be more than an irritant. That could be crippling speed. And so, as the best models—all of them, whether they're domestic or Chinese, whether you are OpenAI or Anthropic or Gemini, whether you're any of the Chinese models—they all moved to a reasoning approach. But that meant more compute was being used during inference. That was a huge advantage for us because we were faster.
Verification, guardrails, and small models checking big models
Andrew Feldman [50:08] And it made the GPU slowness stack up, and it made our speed have an even bigger advantage. And so this was a huge boon for us. We think this is going to stay as a fundamental essence of the way these models are run right now.
Matt Turck [50:24] And you also wrote about verification and whether what was bottlenecking agents was whether the reasoning was good enough, whether they were smart enough, or whether there was a verification problem.
Andrew Feldman [51:01] Sure. I think the verification problem is a little bit like the guardrail problem. And what you'd like to do is, after you've written two or three drafts of your answer, you would like to be sure that it wasn't wrong. And that's your verification step. And you can do that maybe with a different model. You can do that by asking your existing model a similar question in a different way, right? All of these are ways you can pressure-test your result.
Andrew Feldman [51:52] Guardrails can work the same way, right? You want to review, either with a model or with another technique through a scoring mechanism, that this question isn't out of line, that this answer isn't about how to make biological weapons or calling upon information that you are directed to the FBI. And all of that takes compute time. And so whether you're trying to improve through reasoning or whether you're running guardrails, by being faster, you can get results in less time.
Matt Turck [52:04] So what do you think the world is evolving towards? Is it a bunch of smaller models running faster, doing more verification, versus a large model?
Andrew Feldman [52:29] Well, I think those work together. I think your big model produces an answer, then you want to double-check your data. I mean, in the journalism industry, people would write papers that have data checkers, right? Somebody would go and make sure. Back in the world, journalists checked data, right? Being a data checker, right? That was the job. Each claim was checked independently. That's a different model, right? The main model wrote the piece, and then a little model checks some of the answers.
Multimodal AI and the path to video
Andrew Feldman [52:37] And I think that's a very good way to go about it.
Matt Turck [52:44] Where does multimodality for the larger models fall in your world?
Andrew Feldman [53:22] Well, we just announced that we were fastest in the world on one of Google's multimodal models. I think the truth is that there's very little text that doesn't have charts and graphs, right? You must, to understand text, be able to understand illustrations, graphs, charts. And so that's the first and easiest part. And then you ought to be able to create both, right? And then you want to understand images. And I think the new models are very, very good at that.
Andrew Feldman [53:50] Obviously, what follows that is video, because a video is just a collection of images, but that takes an enormous amount of compute right now. And that's one of the reasons it's been set aside by the leading labs. They're so unbelievably computationally intensive, right?
Cerebras’ business model: hardware, cloud, and API
Matt Turck [54:07] Let's talk about the Cerebras business a little bit because you guys are obviously a chipmaker and provider, as we discussed. You're also a cloud provider, data center provider. What are the different parts?
Andrew Feldman [54:37] We make compute, and that compute is optimized for AI. It is the fastest at AI in the world. If you have a data center, we will sell you hardware for deployment in your data center. If you don't have a data center and you'd like to rent it by the month or the year, we have data centers in the U.S., so you can rent our equipment through our data centers and through our cloud. And so that allows us to get AI natives, as well as large enterprises and governments.
Matt Turck [54:48] And what's the proportion in terms of—
Andrew Feldman [55:13] It was 50/50 last year, and I think this last quarter it was maybe 75/25 in favor of hardware sales. I think this year it might be 50/50. And as our OpenAI deal continues to unfold, it'll probably be 30/70, with 30 on-premise deployments of hardware and 70 cloud.
OpenAI’s 750MW inference deal
Matt Turck [55:35] Great. And let's talk about that OpenAI deal, since it's such a major historical, record-making milestone. So it's providing up to 750 megawatts, which is interesting, by the way, as a metric, because you're a chip provider, but this is power. So is that shorthand for—
Why data centers are measured in megawatts
Andrew Feldman [56:13] It's a shorthand. I mean, it turns out right now—and we didn't talk about this because they're sort of in the adjacent supply chain—we went through the shortage of memory. We went through the shortage of a process called CoWoS, 3-nanometer capacity. The other limitation in our industry right now is data center availability. And I mean, that is a limiting factor for everybody. And that's why Anthropic did a huge and sort of very expensive deal with Elon for data center capacity.
Andrew Feldman [56:41] Our deal with OpenAI was because data center capacity is a limiting constraint, measured the way data centers are measured: in megawatts. The deal is 750 megawatts: 250 megawatts in '26 on a multi-year lease, an additional 250 megawatts in '27 on multi-year leases, additional in '28 on multi-year leases.
Matt Turck [56:45] And you're doing the data center for them, or you're providing the chips that go into the data center?
Andrew Feldman [56:52] Data center for them. We're delivering a full cloud solution, so they connect to us via an API.
Matt Turck [56:58] Basically, you mentioned 2026. Do I see immediately that's—do I want to?
Andrew Feldman [57:10] I am looking for data centers. My next meeting is, in fact, with a data center provider. We're doing a lot in Europe right now, a lot in the Nordics.
Matt Turck [57:14] Is that because it's closer to power sources?
Andrew Feldman [57:25] Yes. It is because there is low-cost power, clean, low-cost power, and low-cost cooling.
Matt Turck [57:33] Does it matter, just like in the cloud business, where the data center is located in terms of speed?
AWS Trainium + Cerebras decode
Andrew Feldman [58:01] Yeah, there is an additional latency called transport latency, and that's the speed of light through fiber to get from Helsinki to New York, right? And you have to account for that if your customer's in New York and your data center's in Helsinki. Yeah, I mean, it's usually about two-thirds the speed of light, in case how long it takes. You would like data centers on the same continent, both of ours.
Matt Turck [58:11] That's the OpenAI deal. There was an exciting deal with AWS as well where it's a co-chip solution.
Andrew Feldman [58:23] That's right. It's the disaggregated solution you mentioned before, where their Trainium part is doing the prefill and is doing the parallelizable step.
Matt Turck [58:24] So that's Trainium.
Andrew Feldman [58:42] That's right. Trainium is doing the prefill step, and our chip is doing the decode. And so you get a lot more blisteringly fast tokens. Sweet. It's a good deal for us. It uses their data centers. So these are deployments in the AWS data centers. Yeah.
Matt Turck [58:57] It's fascinating, right? Talking about this industry, how it seems that flexibility is so important. There's, I think, a very good solution. Like, you need chips, you get chips; you need data centers. Everybody's buying from different suppliers to reduce dependency.
Andrew Feldman [59:11] I think that's one of the reasons why we went from being a traditional chip and system provider to also offering data centers, is that what our customers want are fast tokens.
Matt Turck [59:12] Yeah.
Andrew Feldman [59:28] And anything we can do to make the delivery of fast tokens easier—for some of them, that's in their data center; for some of them, it's with an API. Just point your traffic to us, and we'll point the fire hose of fast tokens back.
Fast tokens as a cloud product
Matt Turck [59:54] Yeah. And just to get a sense for where you start and where you stop, you do not, or at least currently, provide the client version. So if I want to run Kimi, like, I know you have incredible stats for Kimi and Gemma in terms of speed, and feel free to mention them, but you don't run those as a service, or do you?
Andrew Feldman [59:55] We do.
Matt Turck [1:00:00] You do? Okay. So you have a service also competing with the Basetens and Fireworks?
Andrew Feldman [1:00:30] Yeah, I think we have an on-demand service where you can come to our site and book a month. I think you can even buy buckets of tokens for Kimi or GLM or some of these models. Many of our customers come there, get excited about it, and then move to a dedicated offering where they take hundreds of machines for a year or two or three or four, once they've proven out the benefit for them in their work.
Andrew Feldman [1:00:43] Often they do A/B tests. Not surprising, people like faster. Okay.
Matt Turck [1:00:45] So, but it's more like a testing—
Andrew Feldman [1:00:47] It's a full environment.
Matt Turck [1:00:47] Okay.
Andrew Feldman [1:00:50] You can go and use it.
Matt Turck [1:00:51] Okay.
Is CUDA still a moat?
Andrew Feldman [1:00:52] Play around.
Matt Turck [1:01:02] Fascinating. But that could become yet another big cloud business. Okay, so you have chips, you have data centers, and you have a cloud business running inference on Cerebras.
Andrew Feldman [1:01:03] So it's going to get interesting.
Matt Turck [1:01:15] Thinking about moats, famously NVIDIA has CUDA as a well-discussed moat. What's your equivalent of—
Andrew Feldman [1:01:17] Well, I don't think CUDA is a moat.
Matt Turck [1:01:18] Correct.
Andrew Feldman [1:02:04] We should talk about that. I think two years ago, every state-of-the-art model was trained in a CUDA flow. And right now, Gemini is trained without CUDA. Anthropic Claude is trained without CUDA. OpenAI is trained with CUDA. So in a one- or two-year period, they lost 70% share of training models that are state of the art, because Gemini is trained on TPUs. It's very good. Anthropic is trained on Trainium, not H100s. And so I think the story of the moat is still present, where the data show the moat is clearly shrinking.
Andrew Feldman [1:02:26] There's no moat in inference. It takes you eight keystrokes to move from a GPU to us in the cloud.
Matt Turck [1:02:31] Eight keystrokes.
Andrew Feldman [1:03:00] That's it, to move your traffic from GPU API to us. And so, obviously, CUDA was enormously important in the creation of our industry, and it allowed the graphics processing unit to be more general than graphics processing. But since 2023, 2024, I think its ability to serve as a durable moat shrunk substantially.
Matt Turck [1:03:12] And you have a whole ecosystem strategy as you think about your moat, to the extent that any moat can be frozen in this industry, right? You build a whole ecosystem.
Andrew Feldman [1:03:24] Yeah. We built an ecosystem. I think our moat comes from the fact that, by virtue of our architecture, we are doing things no one else can do.
Matt Turck [1:03:24] Yes.
How TSMC helped Cerebras build the giant chip
Andrew Feldman [1:03:53] And it's not that they can spend more money or they can't pay more for GPUs. If you want fast, you can't have it. Well, let me shape it. It doesn't work. And so that's where we're building sort of our strength, and that's how we're sort of delivering value to customers.
Matt Turck [1:04:04] How do you think about supply chain? We mentioned supply chain constraints for others, but what are your supply chain constraints? You're all TSMC?
Andrew Feldman [1:04:32] We are TSMC. We have very close collaboration. They were investors in us. They've been exceptional partners. I'll tell you an unusual story. In 2017, we showed up, and we were about 30 guys total, and we showed up in August. Horrible time in Taiwan, right? Don't go to Taiwan in August. The heat is brutal. Not that it's so nice here today. It's only 90 here.
Matt Turck [1:04:33] You're embarrassed to mention.
Andrew Feldman [1:05:06] Yeah. But also the humidity, and you've got to put on a suit. And we met with the leadership of TSMC. We said, "We're like your little pipsqueak company. We believe we can solve a problem that nobody solved in history, and here's how we would modify the way you make chips to make this possible." They thought about it, and they said, "We agree. Let's do it." In the meeting?
Matt Turck [1:05:07] In the meeting.
Andrew Feldman [1:05:12] In the meeting. It wasn't, "Go away for a month." In the meeting.
Matt Turck [1:05:16] Was that a prepared mind, or were they just exceptionally fast on their feet?
Andrew Feldman [1:05:57] First, the salesperson had gathered the decision-makers. Second, our proposal was sort of really good at allowing them to use what they were good at, and it didn't require them to change a huge amount, but it did require them to make real changes. And I think they saw this as sufficiently bold that they would learn as they did it. And they also knew that AI was better on big chips. And so, the combination of prepared mind, willingness to take some risk, bold thinking from a very large company.
Matt Turck [1:06:10] Fascinating.
Andrew Feldman [1:06:17] It is fascinating. I mean, that's how big companies win, right? And how rare is that? It was extraordinary.
Matt Turck [1:06:24] And what happened next? Like, how long does it take between a decision in a meeting like that, which sounds exceptionally fast, to—
Andrew Feldman [1:06:39] Two years. Chipmaking is a long, hard process, and most of the time your first chip isn't a winner. And there are lots of startups now, some of them with really smart guys.
Matt Turck [1:06:40] Yeah.
Andrew Feldman [1:07:03] But their first chip will not be a winner. The TPU—Google had some of the best guys in the industry. First chip wasn't a winner, nor the second, nor the third. Fourth was really good. Now they're on their eighth, and it's a really good chip. The Annapurna team at AWS, first chip wasn't great. Second, third chip really good. It takes time. And so we built a chip, we delivered it, and this gets to an earlier question you asked: we solved a problem that nobody in the history of compute had solved, and we delivered it in 2020, and nobody cared.
Matt Turck [1:07:20] Nobody cared.
Why nobody cared in 2020
Andrew Feldman [1:07:41] Nobody bought any, and nobody cared, which—we were like, "Oh my God." Everybody said we were crazy. It would never work. And now it works, and nobody wants to sell it. It had its hardest. So then we built the next one. I mean, the first one we probably sold 20 or 50.
Matt Turck [1:07:45] And nobody wanted it because the market was not ready, or because the product was not good enough?
Andrew Feldman [1:08:10] Nobody wanted it because AI was a hobby. It was, right? And who cares if your hobby's really fast? You care about fast when it's in production. You care about fast when you use it every day. And so we built another one, and that one we sold 300 or 500. Yeah. And we built the third one, and we sold tens of thousands.
Matt Turck [1:08:11] Right.
Andrew Feldman [1:08:12] Amazing.
Matt Turck [1:08:12] Yeah.
Andrew Feldman [1:08:14] Isn't that interesting?
Why chip supply chains are hard to diversify
Matt Turck [1:08:23] And so, going back to supply chain, do you need to think about onshoring, diversification?
Andrew Feldman [1:08:48] So it's very hard to diversify away from TSMC. Chips are so hard. And you actually, when you design a chip, part of the design is for the rules of that factory, right? So you can't take your design from TSMC and go to somebody else because a huge amount of the work has been to be sure your design is within their rules. And so even in, I think, only one or two exceptions in history, each chip generation goes to one fab.
Andrew Feldman [1:09:20] So that's—we're going to be with TSMC for our next generation as well. I think some of them, we have a supply chain that is built in many parts, but we bring the chips back from TSMC to the U.S. We package in the U.S. and we assemble in the U.S. We do our manufacturing in the U.S., and then we ship from the U.S. I think when you're growing as fast as we are, right, there are a whole range of garden-variety supply chain challenges.
Andrew Feldman [1:09:53] A vendor screws up a batch. It gets stuck in customs. The number of ways that things can go wrong in the supply chain is unbelievable. But we manage these every day, and we're increasing our manufacturing throughput exponentially. And so it's really that part of the business in clover. Wow.
Why today’s AI models will be the worst you ever use
Matt Turck [1:10:09] Incredible. So maybe to zoom out as a last question, what's your best guess about where all of this is going? Obviously, who knows in AI in the next few years, but like in the next—well, we know some things.
What fast AI could do to SaaS
Andrew Feldman [1:10:40] GPT-6 will be the worst model you ever use. And whatever you think is cool about it right now is going to be boring and backwards in six months. And that is so exciting. And I think I watch the way our young engineers use it, and it's very different than the way I'm using it. And it's sort of a fun time where you can learn from your young team members that they're using AI very differently. I think the business of dashboarding and the business of IT, just the damage AI is doing to SaaS is, I think, unrepairable.
Andrew Feldman [1:11:26] For all these, you can ask your AI, "Build me a tool like Salesforce." Thirty seconds later, you have a working tool that is just unbelievable. And all the things that were difficult because they cut across your internal organizational silos, right? One of the things that's really hard is if you want to know, for your top-performing people, when was the last time they got a stock option refresh and how much holding power is left. So, how much unvested stock they have left at today's price.
Andrew Feldman [1:11:36] It was like five systems. You're in Workday, you're in your stock.
Matt Turck [1:11:41] Yeah.
Andrew Feldman [1:11:46] You're in your comp, you're in Carta. None of those can handle, and that's what a CEO wants.
Matt Turck [1:11:47] Yeah.
Andrew Feldman [1:11:48] How much holding power is for my top guys?
Matt Turck [1:11:49] Yeah.
Andrew Feldman [1:12:02] And I used to have little tools I wrote for this and, oh God, I've got a little app that I had. And Spark, right? Swarm.
Matt Turck [1:12:17] Here you go. Well, what a story. It's just incredible to hear all of this from you. What a journey. And what an exciting future. So, thank you very much. I learned a lot, and this was terrific. Thank you, Andrew.
Andrew Feldman [1:12:20] Thank you for having me on your show. I really appreciate it.
Matt Turck [1:12:41] Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.
Andrew Feldman [1:12:42] Bye.