OpenAI’s Compute Chief: We Can’t Build Fast Enough | Sachin Katti

The MAD Podcast with Matt Turck · with Sachin Katti, Head of Industrial Compute, OpenAI

Sachin Katti is the Head of Industrial Compute at OpenAI. We cover why OpenAI sees failing to build enough compute as a bigger risk than overbuilding, why its Jalapeño chips optimize tokens per watt for known model workloads, and why new data centers require investment in generation, transmission, and substations.

Watch on YouTube

Chapters

  1. 1:44 — Is this the biggest infrastructure buildout in history?
  2. 3:41 — Why OpenAI is building a new industrial muscle
  3. 4:54 — What an AI data center actually is
  4. 5:27 — “Factories turning electrons into tokens”
  5. 6:35 — Why AI data centers need liquid cooling everywhere
  6. 8:10 — The power problem: grids, generation, transmission, substations
  7. 10:43 — Behind-the-meter power and gas turbines
  8. 11:02 — Why nuclear “can’t come soon enough”
  9. 11:49 — Jalapeño: why OpenAI is designing its own AI chips
  10. 13:38 — Why inference may now dominate AI compute
  11. 14:58 — Is OpenAI overbuilding compute?
  12. 16:47 — Why OpenAI thinks the bigger risk is not building fast enough
  13. 17:55 — Communities, jobs, water, and the local data-center debate
  14. 21:16 — How OpenAI chooses data-center sites
  15. 22:25 — What “industrial compute” means inside OpenAI
  16. 25:59 — Sachin’s path: Stanford, startups, Intel, OpenAI
  17. 28:05 — OpenAI’s compute portfolio: Microsoft, hyperscalers, neoclouds
  18. 29:37 — Stargate explained
  19. 31:21 — Abilene, Oracle, and the next wave of AI data centers
  20. 32:48 — How massive AI compute gets financed
  21. 34:05 — How OpenAI designed Jalapeño so quickly
  22. 35:59 — AI is starting to help design AI chips
  23. 36:20 — MRC: the networking problem behind 100,000 GPUs
  24. 38:47 — Bottlenecks: transformers, turbines, electricians, supply chains
  25. 40:29 — Guaranteed capacity: intelligence as a supply unit
  26. 42:08 — Will AI data centers move to space?

Transcript

Is this the biggest infrastructure buildout in history?

Matt Turck [1:42] All right, Sachin, welcome. Excited to do this. We are recording this on the sidelines of the RAISE conference in Paris. So thank you for braving the heat. It's another heat wave here.

Sachin Katti [1:44] Thank you. Great to be here.

Matt Turck [2:12] To start, some people describe what's currently happening in the world of compute data centers as the largest infrastructure buildout in history, bigger than the highway, bigger than the railroads. And I'm curious, one, if you agree, and two, what it feels like from the inside. Like, how do you view what you're currently building at OpenAI?

Sachin Katti [2:35] It definitely feels like one of the largest things humanity has ever built, effectively. Definitely bigger than many of the things that I've heard of. I'm not old enough to have experienced the highway buildout, but it feels exactly like what it sounds like. I'm in the belly of the beast, so to speak. Every day, we are making decisions on compute that historically, from my previous role, for example, at Intel, we'd probably take months to make, given the magnitude of those decisions.

Sachin Katti [3:04] But the demand is so insatiable, and it is growing so rapidly that we have to move very quickly. So it's an intense time, but it's probably the most exciting thing an engineer would want to be part of.

Matt Turck [3:15] Yeah. And I read somewhere that OpenAI was planning on spending about $50 billion in compute this year. Is that still the rough number directionally?

Sachin Katti [3:17] Directionally, that sounds about right.

Matt Turck [3:30] And the whole industry itself was going to be $700 billion in compute spend this year as well. So, like, insane numbers, it seems. Yeah.

Why OpenAI is building a new industrial muscle

Sachin Katti [3:41] And it's probably continuing to grow, right? And a lot of build happening. So a lot of that is also going to translate to compute usage from people like us in a year or two.

Matt Turck [4:07] Is the right way to think about this that for OpenAI, it's a bit of a new world, right? So obviously not quite a pivot, because obviously all AI research is going full speed ahead, but like building a whole new business within the company. Is that fair? Is that how people think about it? Because obviously building models is one thing; building data centers, it's a whole different world.

Sachin Katti [4:31] Yeah, I mean, I think OpenAI always has had a fundamental belief that compute is at the foundation of everything, right? Compute is the foundation for intelligence. And the way we keep continuing to scale intelligence and distribute intelligence is by having compute. And so that has never been different. That's always been the belief. I think what's becoming clear is, to build the kind of compute we need, and at this scale, we have to not just rely on getting compute from our partners; we increasingly have to take a much more active role in building and getting that compute that we need.

Matt Turck [4:48] Hmm.

Sachin Katti [4:53] So it does absolutely feel like a new muscle that we are building in the company.

What an AI data center actually is

Matt Turck [5:20] And maybe to anchor the conversation from the beginning, it would actually be very helpful to talk about what a data center is in reality. So I think everybody knows that data centers are being built, but gun to one's head, I'm not sure that everybody could say, well, what is actually being built? Because we've been building data centers as an industry for cloud for decades at this point. So what is fundamentally different and new about the data centers that we're building for AI today?

“Factories turning electrons into tokens”

Sachin Katti [5:57] I think the biggest probably is the scale, right? So we are essentially building large supercomputers as we think about AI. And as we build intelligence and deliver intelligence, and models become more capable, we use it for more and more complex tasks. We need bigger and bigger computers, effectively. And so I think the best way to visualize data centers is giant factories, right, that are turning electrons into tokens. It's a popular phrase nowadays, but it actually has a lot of ring of truth to it.

Why AI data centers need liquid cooling everywhere

Sachin Katti [6:35] So how do we take power, how do we take those electrons and actually use them to power chips that effectively are delivering intelligence? And the way I visualize it is large football fields, liquid-cooled because these chips run really hot. The temperatures on these chips are very, very high. And so you cool them with liquids; you can't cool them with air. So a lot of liquid-cooled, basically refrigerators, effectively, that are sitting alongside the building.

Matt Turck [6:44] And on that topic, while we're at it, does the cooling happen at the data center level, or does it happen at the chip level, or both?

Sachin Katti [6:44] Both.

Matt Turck [6:45] Both.

Sachin Katti [7:07] Right, so you need to cool the data halls, but you also need to cool the chips individually, because it's not going to be enough to do one or the other. And you also have to cool the things that connect chips, right? And so that's why you need cooling pretty much everywhere nowadays. Even the cables, the transformers that distribute the power, become too hot, so they also need to be cooled.

Matt Turck [7:21] So everything that processes energy produces heat. And is the cooling technology that is being used something that's well understood and is just getting deployed, or are there fundamental new things happening in cooling right now?

Sachin Katti [7:49] I think liquid cooling has been around for some time but has never been deployed at this scale. And so the innovation is more around how to make it reliable, how to make it cheaper, more scalable, right? So there's a lot of innovation around that. There's also a lot of new innovation, new kinds of liquids, new kinds of materials that can absorb heat better, because anything that can improve the efficiency of heat transfer is very important for data centers.

Matt Turck [7:50] Hmm.

The power problem: grids, generation, transmission, substations

Sachin Katti [8:10] So we can then run the chips hotter, right? And there is a direct correlation between running a chip hotter and how powerful the compute is. So the hotter the chip, the more memory bandwidth you get, the more FLOPs you get. And so there's a strong payoff. If you can cool well, that also means you can produce more intelligence.

Matt Turck [8:29] All right, so gigantic factories, lots of cooling. The other part that seems to be very critical to any discussion is power and energy. So how does that work, starting at a high level? Do you connect to the grid? Do you have your own power generation?

Sachin Katti [8:53] I think in the early days, we all connected to the grid, and we still all would want to connect to the grid. At this point, we are beginning to hit, and we are investing in generation infrastructure for the grid, transmission infrastructure for the grid. So whenever we build a data center anywhere, we make it a hard commitment that we are not taking power away from the grid. In fact, we are investing in the grid to generate new power so that we can consume it for data centers.

Matt Turck [9:01] What does that mean practically?

Sachin Katti [9:22] You have a grid somewhere. It has a certain generation and distribution capability, a certain number of megawatts. Obviously, a data center shows up. If there was spare capacity, then of course the data center can use it. But if there isn't spare capacity, then we have to add new gas or solar or hydro generation infrastructure to the grid. So we are investing and funding that build-out, and then you have to build transmission lines, invest in transformers, substations to distribute that power.

Sachin Katti [10:02] So wherever we are building data centers, we are funding the development of all of that infrastructure. And so that's one of the things that we do want to emphasize, which is this is infrastructure that would otherwise not have been funded if not for these data centers. And one of the side benefits of this big data center build-out is the grid infrastructure of America, and the whole world for that matter, is getting upgraded very quickly. And so that's the power piece. So whenever we can do that, we do that and we consume power from the grid, but we are good citizens of the grid because we are improving the infrastructure for everyone.

Matt Turck [10:14] Yeah.

Behind-the-meter power and gas turbines

Sachin Katti [10:44] Not just for data centers, but also for households. In some places, we are beginning to hit the limits of how much grid power we can build and consume. And so, everyone's looking at behind-the-meter. We are also doing some behind-the-meter generation, where we would have on-site power generation and distribution capability that does not come from the grid, where, in fact, the data center becomes effectively self-sufficient in terms of power.

Matt Turck [10:46] That's gas turbines? What is it?

Why nuclear “can’t come soon enough”

Sachin Katti [11:03] Today it's gas turbines, especially in the U.S., because that's the most dense, transportable form of energy, and also the one that is quite widely available in the U.S. But there, we are bottlenecked by the supply chain.

Matt Turck [11:20] Do you think the nuclear conversation is interesting? We're recording this in France, which has a bunch of nuclear power generation. Nuclear seems to have come back to the discussion in the U.S. as well. Is that something that you think about and you think is interesting? Absolutely.

Sachin Katti [11:47] It can't come soon enough. I think it is the densest form of energy we can all produce and consume, and it's also clean. So I think definitely would be a good source of massive, scalable energy for data centers. Obviously, outside of France, the rest of the world has a lot of catching up to do in building this infrastructure, but I think it's going to play a very important role.

Jalapeño: why OpenAI is designing its own AI chips

Matt Turck [12:17] Okay, so that's a great introduction on data centers. The other interesting bit of news that you guys recently had is Jalapeño. So now for OpenAI, in addition to being in the application business, consumer and enterprise, and then being in the model AI research business, and then the compute and data center business, it seems that OpenAI is in the chip business, if that's fair. So, like, completely full stack. But I'm curious, and we'll go into some details about Jalapeño later, but in terms of overall strategy, where does that fit?

Sachin Katti [12:54] As we begin to serve a pretty big fraction of the world's population, AI usage is exploding. Inference is obviously becoming a big fraction of our workload, and it's consuming a lot of compute. And one of the other realizations is because we know what the workload exactly is, what the model we want to run is, we can co-design the hardware to be super efficient in delivering those models. And so the strategic thesis we have at Jalapeño is: how do we take advantage of knowing what the end workload is, what the model itself is, and design chips that are very efficient in serving those models?

Why inference may now dominate AI compute

Sachin Katti [13:39] So it really allows us to drive an efficiency advantage, drive more tokens per watt. So the key metric that Jalapeño is optimizing is maximizing the number of tokens you can produce per watt. And because the world is constrained by power today, the more tokens you can produce for the same amount of power, the better it is for everyone. So we look at it as a very critical ingredient in scaling how we deliver intelligence to the world. Great.

Matt Turck [14:15] So we'll come back to helping in a second. But you just mentioned inference, and it's such an interesting evolution as well. Without commenting on necessarily what's going on at OpenAI specifically, is inference equally big or much bigger than training these days in terms of usage of compute? Have we shifted from those very heavy pre-training runs as the major use case for compute to now just inference being the majority?

Sachin Katti [14:41] No, inference is big, perhaps even the majority of compute. And I think we don't like to make a distinction between training and inference because a lot of training is now inference. So when we train a new model, we are generating synthetic data, for example. That's inference. When we train a new model, we are doing post-training.

Matt Turck [14:41] Yeah.

Is OpenAI overbuilding compute?

Sachin Katti [14:58] And that's inference. When you train a model, you're doing test-time compute. That's all inference. So when we say training, a lot of the compute actually is inference, even in that phase of the work. So inference is a fundamental building block.

Matt Turck [15:31] Obviously, I cannot resist asking the inevitable question around the potential risk of overbuilding, given the lag between demand and usage and how long it takes to build a data center. I think you mentioned somewhere that you were deliberately very paranoid about the problems ahead in the next three years, very paranoid about the surprises ahead, which sounds like a very healthy approach. So how do you think about that? Is there any way to mitigate that, or is it just ultimately a deep belief that this is the future and we should all just go, go, go?

Sachin Katti [16:12] We have deep conviction in scaling, and history has borne us out. So effectively, our revenue, for example, has tracked the compute. We tripled compute, and we tripled revenue. And we believe that continues to be true. Demand far outstrips compute supply today. So anything we can bring online, we consume immediately. So there's no compute that is going to waste, for us at least. So I think that conviction has not changed whatsoever. And if anything, we are seeing that scaling laws on research and training continue to hold, and potentially the pace at which we are doing research is accelerating because of AI itself.

Why OpenAI thinks the bigger risk is not building fast enough

Sachin Katti [16:51] So, AI is doing a lot of AI research now. And so, one of the subtle implications of that is previously our researchers used to run experiments, and they needed compute to run experiments, but the number of experiments they could run was limited by the number of human researchers they had, which is a scarce resource in the world. There's not a lot of people who can do AI research. Now, if AI itself can do AI research, the number of experiments we can run explodes.

Sachin Katti [17:08] And therefore, the amount of compute you need for research also explodes. So, we don't see a world where we will have unused compute for the foreseeable future.

Matt Turck [17:09] Right.

Sachin Katti [17:17] When I was referring to surprises, my worry is more on the downside of we are not able to actually build all the compute we want.

Matt Turck [17:20] The risk is the other way.

Sachin Katti [17:45] It's the other way for us, right? Because that is consistently the case. Any time we have thought we have enough compute, we can slow down, it has always negatively surprised us, like, oh shit, we should not have slowed down, right? And so our biggest worry is that still. And at the scale at which we are planning to get compute and build compute, the physical world does not move that fast, right? Physical supply chains, factories don't move that fast, cannot add capacity that fast.

Communities, jobs, water, and the local data-center debate

Sachin Katti [17:55] So for us, the surprise is more in that direction than the other direction.

Matt Turck [18:28] Okay, fascinating. You alluded to communities a minute ago, and obviously that's a key debate. So, curious about your perspective on a spectrum where, on the one hand, at one extreme, you'd say, well, the AI industry and compute industry has a PR problem, and there's no problem. It's just that we cannot explain it well enough. To the other extreme, actually, those communities have a point. What do you think the reality is?

Sachin Katti [19:00] I think anytime there's new technology which is as revolutionary as this technology is, there is always disruption that's going to happen. But inevitably, we have learned this over history that this always leads to better outcomes for society, right? And so, how do we draw a line from where we are today to that outcome, right, and explain to the world why this is the trajectory we all need to be on? I mean, it's our responsibility to do that. On the communities front, there's a little bit of a local-versus-global issue, right?

Sachin Katti [19:34] On the communities, I think data centers are, even today, a net positive to every community because we are building these data centers in rural areas of America, for example, right, where there's nothing else that is being built on this scale. So, we show up in rural Texas, we build a data center, that produces new property tax receipts for the community, that funds schools, that funds hospitals. We show up and we invest in new grid infrastructure, which otherwise would never happen because there's no demand.

Sachin Katti [19:47] So there's a modernized grid that that area can enjoy. Basically, we produce jobs.

Matt Turck [19:48] Nice.

Sachin Katti [20:19] And so I think one of the things we are investing a lot in is explaining the local benefits every time we build a data center somewhere and making sure that it is well understood, the kind of upside that this has. And data centers, once they're built, are essentially very clean citizens, right? They don't produce any gases or toxic chemicals or anything like that, right? They're self-contained. They just produce intelligence.

Matt Turck [20:33] The typical question that comes up is water, and I think that's been debunked quite a bit by research, but maybe give us just color on how you all think about the water question.

Sachin Katti [20:58] These are liquid-cooled, and the liquid is recycled. So, the water consumption of a data center is shockingly small relative to household water consumption. So, I think as you put it, it's been debunked. It's a misperception that data centers consume a lot of water. If anything, they consume so little water for what they do, and all of that water is recycled. So we don't net consume new water. Once we get to a particular point, the water just gets recycled as we use liquid cooling.

Matt Turck [21:14] Yeah. So all those stories about, like, groundwater, they just don't make sense because the water at a data center happens in, like, a contained circuit, right?

How OpenAI chooses data-center sites

Sachin Katti [21:16] It's a closed loop. Yeah, it's a closed loop. It's a closed loop.

Matt Turck [21:31] And by the way, you mentioned Texas and rural areas. Since we talked about data centers at the beginning of this conversation, why do OpenAI and other companies pick rural areas? Like, how do you select a site for a data center?

Sachin Katti [22:00] So many factors. So one is, of course, land, like plentiful land. Number two would be permitting. Like, can we build these things? And we want to build these things such that they are not affecting any neighborhoods, right? So land that is somewhat removed is the ideal candidate. Of course, access to power, right? So a strong grid, strong gas availability, all of those are important factors. And then four is labor, right? So how quickly can you build these things?

What “industrial compute” means inside OpenAI

Sachin Katti [22:27] So availability of labor, construction labor, qualified electricians, plumbers, all this has to be available. So all of those factors go into every single site selection decision. And I know obviously Texas has been popular because it fits a lot of these criteria, but it's not the only state. I mean, we have data centers all around the country. Okay, great.

Matt Turck [23:00] We're going to go into all of this in more detail, but let's talk about you a little bit and your journey. So you're the Head of Industrial Compute at OpenAI, which, by the way, to the beginning of this conversation, just the title Industrial Compute is such a perfect title for the moment we're in. But what does that mean? What is the role, and how is this whole effort organized within OpenAI, to the extent you can talk about it?

Sachin Katti [23:28] Yeah, I think of it as my role and our team's role, rather, as how do we bring compute online at industrial scale, right? That's effectively what we're doing. And that's the entire lifecycle. So, how do we find the ingredients that go into compute? Land, power, shells, chips. How do we finance them, right? Because these are massive dollars. And so, how do we make sure that we finance the grid infrastructure? How do you finance the construction of the compute shells?

Sachin Katti [24:00] How do you finance the chips? Then it's about how do you operationalize all of this? So how do you actually make sure these things happen on time, they stay up? How do you operationalize all of this infrastructure? So it's that entire lifecycle. And then, of course, how do you actually use the compute? So a big part of my role is capacity allocation inside OpenAI. So it is always a scarce resource.

Matt Turck [24:02] I'm sure that makes you a very popular guy.

Sachin Katti [24:23] I am not very popular. There is always someone who is unhappy with whatever decisions you make. But yeah, our team provides the input to make the capacity allocation decisions. So we surface what are the different choice points and what are the what-if questions, different allocation choices that we have. So capacity planning, and then, of course, using that to forecast how much capacity we need where, because it's not just more compute, it's also where, what kind, what shape, what chip, what workload you want to run there.

Sachin Katti [24:47] So this team figures out what should be the forecasting and planning, and that closes the loop.

Matt Turck [24:47] Right.

Sachin Katti [24:53] So that informs where do I go find the next chunk of land and power and chips to put into.

Matt Turck [25:11] And again, without going into anything confidential, although I guess when you guys go public, all of this will soon be public, but is that thousands of people at this stage? I mean, is that multiple different teams, or do you guys kind of outsource a bunch of things and work with a bunch of contractors?

Sachin Katti [25:38] It's a portfolio approach, right? So we are never going to be in a world where we outsource everything or build everything ourselves, right? It's always going to be a mix because that's the reasonable thing to do, right? So you don't want to put your eggs all in one basket. So we will have hyperscalers probably providing a big chunk of our compute, a majority of our compute. We will have new clouds as parts of our portfolio. We will be partnering with design-build firms that can build the compute that we need.

Sachin’s path: Stanford, startups, Intel, OpenAI

Sachin Katti [25:59] And of course, we may build some of it ourselves. And so we are always going to have a portfolio approach because, at the scale which we need, we will need to tap into all sources of compute. We can't just rely on one particular mechanism.

Matt Turck [26:11] And your background before all of this, so you are both a professor at Stanford and an entrepreneur or a founder, or mostly an academic? Like, just walk us through.

Sachin Katti [26:24] A bit of all of the above. But yes, at my heart, I'm an academic. So I've been a professor at Stanford since 2010. Recently—

Matt Turck [26:26] What do you focus on there?

Sachin Katti [26:29] I was faculty in computer science and electrical engineering.

Matt Turck [26:31] With, like, a particular interest?

Sachin Katti [27:02] Yeah, my area of research was networking. So I did build networks, both mobile wireless networks other than cellular and data center networks, actually. But three or four years ago, while I was at Stanford, I did a couple of startups. The last startup got acquired by VMware. And that's how I got to know Pat Gelsinger, Intel's CEO. And that's how I ended up at Intel. Most recently, before coming to OpenAI, I was Intel's CTO. And so I've kind of seen all the different things: academia, startups, corporate at Intel, and then, of course, a mix of all of the above at OpenAI because we have a research lab, a startup, and a fast-growing company all mixed into one at OpenAI.

Sachin Katti [27:20] Yes.

Matt Turck [27:25] Is that what you said? Why did you say yes to the job when the job came up?

Sachin Katti [27:57] It's actually what I just said. That makes this so unique, and it's hard to find anywhere, right? Because you always have to choose. But having a world-class research environment coupled with the hardest technical problems—like, we are building the largest compute in the world—and so there are a lot of new problems that we need to solve, but also a fast-growing business. I like problems that sit at the intersection of business, technology, and strategy. And so this is, like, a very unique time in history and a unique role, which is very attractive, obviously.

OpenAI’s compute portfolio: Microsoft, hyperscalers, neoclouds

Matt Turck [28:35] Very cool. Going into a bit more specifics about OpenAI's compute strategy, maybe let's summarize what you guys currently have. I think there's some Microsoft. You did this big $20 billion deal with CoreWeave. There's a bunch of things going on, Stargate. Maybe just give us the lay of the land of what you currently have, and then we'll talk about what you're building next.

Sachin Katti [29:08] We have compute from effectively many sources. So Microsoft, obviously, is a big partner, important partner. We also have compute, as we have announced, from AWS and Google. So we have compute from all of the hyperscalers, effectively. We also have compute from CoreWeave, for example, so a neocloud. And then, of course, compute that chip partners are supplying now, like NVIDIA itself directly providing compute for us. So I think that's the mix roughly today. As we go forward, obviously, we'll be building on all of these relationships, but also looking at more options where we design the compute, the data center itself, ourselves, or potentially even build a data center ourselves.

Matt Turck [29:24] Yep.

Stargate explained

Sachin Katti [29:38] So, all of those are ways of scaling the amount of compute that we have. So, I think coming back to my earlier answer, the answer always will probably be trite, but it's all of the above.

Matt Turck [30:07] Yeah, that's helpful. And obviously, diversification makes all the sense in the world given the scarcity. And then it seems that Stargate has evolved from what was going to be a joint venture with Oracle and SoftBank to what now seems like it's more like an umbrella term for the compute strategy these days. Is that a fair way to describe it?

Sachin Katti [30:35] Yeah. I mean, we look at Stargate as our compute strategy, and it is varying degrees of us designing or building the compute ourselves. For example, with Oracle, close partnership, we help them design, we help them with how to operate AI compute, which is, in fact, a very new kind of compute for us. We work with SB Energy, which is public. We basically have co-designed the warm shell with them, and they're executing on that warm shell, and we will be kind of figuring out how to operate our chips in these data centers ourselves, even the new chips.

Sachin Katti [31:11] So, Stargate to us is that umbrella strategy across all of these different things. And think of it as an evolution that we'll continuously be on, because it's never going to be that tomorrow we wake up and do only one kind of way of building compute.

Matt Turck [31:12] Yep.

Sachin Katti [31:20] I think Stargate to us is a continuous way of learning how to scale compute, and we'll be adding on more and more capabilities.

Abilene, Oracle, and the next wave of AI data centers

Matt Turck [31:28] And as part of that, there are data centers being built, right? Like Abilene, Texas.

Sachin Katti [31:29] Yes.

Matt Turck [31:36] So maybe walk us through that. What is currently being built, for people to have some situational awareness?

Sachin Katti [31:54] We obviously have a big partnership with Oracle. That's the Abilene data center. That's where, for example, we are training our newest models. So, very excited about that. That's a very big GB200 Blackwell cluster for our needs.

Matt Turck [31:56] So that's up and running?

Sachin Katti [32:12] It's up and running. It's been used for training the last two models, more actually. So we're super thrilled about that, and you're seeing the results, right? You're seeing how quickly the models are becoming more capable. It's because of these kinds of compute.

Matt Turck [32:19] Other data centers are currently being built that are good at the right—

How massive AI compute gets financed

Sachin Katti [32:49] Yes. So Oracle is building a number of data centers, all of which are public, across Michigan and Texas and other places. So these are coming online in the next couple of years as they get built, and we put whatever chips at that time are the latest for training. But these are really meant to be very big clusters that allow us to do both training, but also product inference-kind-of compute.

Matt Turck [32:54] Right. But the way those deals are structured, Oracle is the prime building those?

Sachin Katti [32:56] Oracle is the cloud.

Matt Turck [32:57] Is the cloud.

Sachin Katti [32:58] And we're consuming compute from Oracle.

Matt Turck [32:59] And you're the tenant.

Sachin Katti [32:59] Yeah.

Matt Turck [33:00] You're the core tenant on those. Correct.

Sachin Katti [33:01] Interesting.

Matt Turck [33:32] For the stuff that you're building, how does the financing strategy work? I mean, obviously you all raised, what was it, $122 billion? Was it the number recently? So there's no shortage of cash, although, given all those expenses, I don't know. But the CoreWeaves of the world are famous for being very strategic users of debt. Is that part of the financing strategy as well? How do you all think about it?

Sachin Katti [33:55] I mean, with all of this compute that we have today, we have amazing partners who actually are handling that for us. So Microsoft, Google, Amazon, Oracle—we are the offtake. So we are the tenants, as you put it. So we commit to consuming that compute, to buying that compute, whether it's on-prem.

How OpenAI designed Jalapeño so quickly

Matt Turck [34:05] Oh, so across the board, like everything, because you're building stuff as well. So the partners are building.

Sachin Katti [34:06] Partners are building.

Matt Turck [34:11] And you're always, across the board, the tenant, not the owner.

Sachin Katti [34:12] Correct.

Matt Turck [34:36] Okay, so therefore financing is outsourced to partners. Okay, we talked about Jalapeño. Let's go into a bit more detail there, because in particular, it seems that you guys went incredibly quickly designing it. I think I read somewhere, nine months from design to tape-out. So maybe walk us through that, and what was the reason it went so quickly?

Sachin Katti [35:06] Yes, it was incredibly quick. Nine months is very, very fast, probably the fastest I've seen in my career. I think several reasons. One is it's a team, it's a strong team. Many of the team have designed TPU chips at Google in the past, so a very well-experienced team. We have a great partner in Broadcom that has a very strong track record of delivering XPU ASICs. So I think that strong partnership with Broadcom in making this happen. Third, I think perhaps an OpenAI-unique point.

Sachin Katti [35:39] In most chip companies, or when you design chips, you don't know what you're designing it for because you're a vendor. The customer who eventually runs the workload is someone different. There's a unique advantage here of us knowing what the future models might look like, and therefore being able to short-circuit a lot of the decisions you need to make, design decisions you need to make on the chip side. So that's super helpful. And finally, increasingly, AI itself helping design and optimize the chip.

AI is starting to help design AI chips

Sachin Katti [36:00] That is usually the one that takes the longest time, because you're basically limited by how much human time there is to process all this data and run the experiments. And we can do a lot of those iterations much faster using AI. Yeah.

Matt Turck [36:03] So AI is building its own chips now?

MRC: the networking problem behind 100,000 GPUs

Sachin Katti [36:20] Yes. I think that world is not very far. I mean, AI right now is assisting in chip design, but we do believe that the world of recursion is not that far, where AI will design the systems it needs to train and run the next generation of AI.

Matt Turck [36:21] And including chips?

Sachin Katti [36:22] Including chips.

Matt Turck [36:35] Including chips. You also released, a few weeks ago, MRC, which is a networking protocol. Walk us through that. What is it, and why is it a big deal?

Sachin Katti [37:05] It's a new networking protocol, routing technology, if you will, to scale these really large cluster fabrics. So imagine you have 100,000 GPUs. They need to be connected together. And when you're doing large training runs, they're constantly communicating with each other because these models are so large that the processing of the models is happening over the entire 100,000-GPU cluster, for example. And you can imagine the number of links and switches and NIC cards that need to be there to connect all of these chips together.

Sachin Katti [37:47] At this scale, failures are common, right? Happens all the time. You can't really even enumerate all the ways things could fail. So the strategy behind MRC is: how do you design algorithms and protocols that can gracefully mask all these failures and make sure that the training workload does not get impacted? The network is an abstracted system that the training job does not have to worry about. It's always going to be there. It's always going to find a path, even if a link fails.

Matt Turck [37:52] Hmm.

Sachin Katti [38:13] So it was all about reliability. It was all about availability. So how do we design protocols that can tame the complexity of such a big cluster and make sure that we don't get stopped because of failures, which are very common in the system? So MRC is like a multipath spraying protocol where you can spray packets over multiple paths. So between any two chips, there are many, many routes to get there, kind of like between any two points in the city, there are many routes.

Bottlenecks: transformers, turbines, electricians, supply chains

Sachin Katti [38:47] So instead of picking just one route, we will send traffic across all of them. And whichever one succeeds, you take that, right? And so that way, even if any one of them fails, it's not a showstopper. So that's kind of the basic intuition behind it. It's obviously a lot more sophisticated than that, but doing this at scale, at that speed, is hard. And so that's why it's quite innovative.

Matt Turck [39:01] What are the bottlenecks that you experience these days? So it seems like the nature of the bottleneck keeps changing in the compute industry. Obviously, people talk a lot about memory these days. Is that one of them, or what else?

Sachin Katti [39:25] I think there are bottlenecks everywhere, to be honest, in the supply chain. So I don't think there's any one, right? We have bottlenecks in the data center building itself around permitting, around availability of gas turbines, transformers. Those industries historically have not added much capacity over the last decade or so.

Matt Turck [39:25] Yeah.

Sachin Katti [39:28] And they've suddenly experienced a demand shock.

Matt Turck [39:28] Yeah.

Sachin Katti [39:35] And it takes years before you can add capacity to produce more turbines and transformers. So they're trying to play catch-up.

Matt Turck [39:46] And to the jobs thing that you mentioned earlier, is that like a shortage of electricians and technical people, trained people that know how to build those things?

Sachin Katti [39:48] Absolutely. Yeah, absolutely.

Matt Turck [39:54] So that's what we should all become as AI replaces knowledge workers, like electricians?

Sachin Katti [40:19] No, I think there is definitely a shortage of electricians, plumbers, all kinds of trades, you name it. So anything we can do to train more folks to be able to do those, they are very well-paying jobs that a lot of us, all of the hyperscalers, all of the labs, would actively hire you for if you had the qualifications. So I think that is definitely a bottleneck. I think it'll become an increasing bottleneck because we are all trying to build more and more, and we have a limited number of these capabilities.

Guaranteed capacity: intelligence as a supply unit

Matt Turck [40:52] Great. Let's talk for a minute about the business side of things. Another thing you launched recently is guaranteed capacity for customers to lock in compute, which is interesting, right? It almost feels like OpenAI is also becoming a utility company, providing compute to others. What's the story behind that and the strategy?

Sachin Katti [41:21] Yeah, so guaranteed capacity is guaranteed tokens, right? So we are effectively saying we will guarantee you a certain dollar's worth of tokens of intelligence. And it makes sense, right? In a world where compute is in short supply, tokens are always going to be at a premium, and there's a shortage of tokens that we can produce given the limited compute that we have. And as this becomes a fundamental input to the enterprise, enterprises are going to need intelligence, more and more of it, to run, right?

Sachin Katti [41:52] And so this is a way for enterprises to gain assurance that the tokens of intelligence that they need will be there for them, and so they don't have to take business risk, right? And so I think it's good business hygiene. If you have a critical supply resource, and for every enterprise, intelligence is kind of the most important supply item, if you will, it makes little sense not to make sure you secure that supply. And so that's the demand we are trying to fulfill here.

Will AI data centers move to space?

Sachin Katti [42:08] I know it's a new concept, like, what does it mean to have guaranteed capacity for intelligence? But I think that's what intelligence is becoming. It is becoming a supply unit for every digital enterprise.

Matt Turck [42:28] So maybe to end on a fun one, and you can answer either with your OpenAI hat on or not: data centers in space. Is that exciting? Is that science fiction? Is that needed? Is that something that people like to talk about just because it's cool, or where do you land?

Sachin Katti [42:53] For the geek in me and the engineer in me, it's definitely exciting. It's one of those things that can be super cool to see, a constellation of satellites producing compute. I actually think it will become feasible, right? As engineering problems go, they can be solved with enough time and investment. Whether it is needed, I think there is room for orbital compute. I don't think it's going to serve all the compute needs, but it's definitely going to be a component in the arsenal.

Sachin Katti [43:28] I think what we are waiting for is: when does the economics of launching satellites change, and when does the economics of the hardware change? Because we need to get to a point where it's cheap to launch the hardware, and if something fails, it's cheap to throw it away, because you can't go up and fix it, unlike on the ground. And so I think that inflection point hopefully happens soon, and at that time it becomes viable.

Matt Turck [43:32] Great. Sachin, it's been wonderful. Thank you so much for spending time with us. Really appreciate it.

Sachin Katti [43:36] Thank you. It's been great to have this chat. Cool.

Matt Turck [43:56] Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you on the next episode.