The GPU Myth: State of AI Compute 2026 | Stephen Balaban
The MAD Podcast with Matt Turck · with Stephen Balaban, Co-founder and CTO, Lambda
Stephen Balaban is the Co-founder and CTO at Lambda. We cover why AI compute is a vertically integrated service rather than a commodity, how H100 price indices can mistake contract mix shifts for falling demand, and why land, power, and data center shells are the industry’s main bottlenecks.
Chapters
- 1:21 — Why GPU compute was never a commodity
- 2:45 — The H100 price index and what it gets wrong
- 4:02 — The real moat: technology or financing?
- 5:57 — Winner-take-all, or room for many neoclouds?
- 6:48 — Are we overbuilding or underbuilding AI compute?
- 9:26 — What if AI gets 10x more compute-efficient?
- 10:44 — The real bottleneck: land, power, and shell
- 11:38 — The backlash against data centers — and the misinformation
- 15:00 — Opening the hood: from photons to tokens
- 17:11 — Extracting more value from the same chip
- 19:26 — Frontier inference and distributed training, explained
- 23:26 — What actually drives compute cost
- 25:21 — Lambda's chip stack and the NVIDIA relationship
- 26:17 — A multi-silicon world? CUDA, CUDNN, and NVIDIA's real moat
- 28:59 — Networking, storage, and the one-click cluster
- 34:46 — Renting vs. owning, and full vertical integration
- 36:24 — How global is Lambda? Does location still matter?
- 38:44 — The financing stack: off-take agreements, SPVs, and credit
- 41:16 — Why a 2023 GPU leases for more today
- 42:36 — A futures market for compute?
- 43:54 — Origin story: facial recognition, Perceptio, and Apple
- 47:03 — The Lambda hat and Dream Scope
- 48:59 — The $60K bet that became a cloud business
- 52:00 — Holding the team together through the hard times
- 54:30 — Bringing on a new CEO; Stephen as CTO
- 57:33 — Matching xAI on high-velocity deployment
- 59:29 — "AI won't write software — it will become the software"
- 1:01:30 — Neural software vs. vibe coding
- 1:04:25 — Do agents change the compute layer?
- 1:06:14 — Self-assembling software inside Lambda
- 1:08:18 — Gigawatt-scale AI factories
- 1:08:57 — One person, one GPU
- 1:12:04 — Hot takes: overrated and underrated in AI
Transcript
Why GPU compute was never a commodity
Matt Turck [1:36] There was a moment in time in Silicon Valley a few years ago, if you had asked most people, they would've said that neo-clouds were going to be a commodity, in particular because GPU compute was going to get commoditized. And if you fast-forward to today, it seems to be exactly the opposite.
Matt Turck [1:50] Both Lambda and several of your competitors seem to be absolutely ripping. So what is it that naysayers got wrong then and continue to get wrong today?
Stephen Balaban [2:28] The big thing is that cloud compute is not a commodity service. It is a very complicated, highly vertically integrated type of service that spans everything from land entitlement, construction, HPC—high-performance computing—design, software virtualization, cloud services on top. And there's a reason why the biggest companies in the world, these multi-trillion-dollar market-cap businesses—whether it's Amazon, Microsoft, Google, or Oracle—are all in the cloud computing business: it's because it's a great business. And so I think that's probably the fundamental thing that was misunderstood, is that, oh, this is somehow a little bit different than a normal cloud service.
The H100 price index and what it gets wrong
Stephen Balaban [2:45] But really what it was, was it's a cloud service designed for the age of AI.
Matt Turck [2:55] But there is some element of commoditization, right? The price of renting a GPU is going down. But what you're saying is that, to some extent, it doesn't matter because it's only one layer of the cake.
Stephen Balaban [3:26] Yeah. So when you look at, for example, I think it's actually worth trying to dig into some of the methodology on an index like the Bloomberg index for H100 rental prices. And what we're actually seeing in the market is that, first of all, there are two different rates. There's a public cloud on-demand rate, and then there's a long-term rental rate. And I think that some of these indices don't properly take that into account.
The real moat: technology or financing?
Stephen Balaban [4:02] Because what we're actually seeing is a very consistent, if not increasing, long-term rental rate, and very consistent and increasing on-demand rental rates. And so what happens is, if the index mix, for example, if the methodology in the index biases towards long-term contracts being a bigger part of the volume, that will look like a decline in the index, when the reality is it's just a decline in the mix that the index is covering.
Matt Turck [4:23] Fascinating. So I'm curious about your thoughts as a key leading player in the neocloud ecosystem about how you see the market evolve. How much of the competitive advantage that you guys are building, and other players are building, is based on technology versus financing?
Stephen Balaban [4:52] There's a few different layers on it, which is, there's a lot of differentiation and work that's being put into, for example, the cloud software orchestration layer, which allows us to, for example, take a very large-scale GPU cluster and partition it up for our customers. So we've got, for example, our One-Click Cluster product that allows us to do that. And that's something that's quite unique in the neocloud space. Most of the other neoclouds either don't have the ability to launch a cluster from their website or max it out at, let's say, 32 GPUs, whereas Lambda's designed a piece of software that allows us to give you anywhere from 16 up to 4,000 GPUs in a web interface.
Stephen Balaban [5:34] And then there's innovation on the data center construction and design side of things, which is also really important, right? Because that's the physical layer underneath the high-performance computing equipment. And we're working on a lot of different ways to dramatically reduce the time it takes to construct and stand up new megawatts. And then there's, as you mentioned, innovation on the finance side of things, where we're coming up with new and unique ways to finance, underwrite, package these large-scale capital projects, really.
Winner-take-all, or room for many neoclouds?
Stephen Balaban [5:57] And so I think innovation's happening on every layer of the stack, and it's a very complex coordination-style business.
Matt Turck [6:06] Yeah. And do you think that ultimately the new cloud ecosystem becomes a winner-take-all, or is there room for multiple very large players?
Stephen Balaban [6:40] No, I think there's absolutely room for multiple very large players, just like the traditional cloud business has shown that there's room for multiple large winners and multiple large players. And I think the fundamental reason for that, kind of going back to what drives market structure, is that generally speaking, when you have an industry that has technology moats and capital formation moats and economic moats, that tends to be oligopolistic in its market structure. When you have markets that have more sort of network-effect moats, those tend to be a little bit more single-winner-take-all.
Are we overbuilding or underbuilding AI compute?
Matt Turck [7:00] What are the various scenarios in your head as you think about the future, about how it all plays out? Are we overbuilding? Are we underbuilding? Nobody knows. How do you think about it?
Stephen Balaban [7:21] Well, I think that we continue to be generally underbuilding. And most people that are sort of in leadership positions at neoclouds or within the market have been recognizing this sort of insatiable amount of demand for large language models to do everything from being an assistant to code generation. You can kind of look back to some of the talks that I've given in the past, around—I kind of called, hey, in a couple of months to years, we're going to be at a point in time where you can put money in and get software out the other end.
Stephen Balaban [7:55] And now, at that point in time when I was predicting that, it was maybe not as widely held of a belief. But now, with, let's say, the release of Opus 4.5, I think it's pretty clear that we have an amazing system that can take in money and output software. And I think the part which makes me feel so confident that there's going to continue to be demand is that we continue to see no end to the scaling laws, which are like the underlying idea that you put more compute in and you get better intelligence levels out of your models as you increase the capacity of the model.
Stephen Balaban [8:45] Train it with more compute, train it with more data, you get more intelligence out. And as long as that continues to hold, I think that we still have in store for us—it's hard to predict exactly when scaling laws might start to reach sort of a diminishing marginal return type of part of the curve. But right now, it's very clear that we're going to continue to see more and more and more capable models. That is kind of expanding the cone of the addressable market.
Stephen Balaban [9:21] Originally, the cone of the addressable market was, all right, this is going to be helpful for customer support. It's a sort of substitute good for Google Search and for other search online. And then now it's like, well, this is a substitute for a lot of software engineering roles or a huge augment to software engineering roles. And so as that cone expands, the total market and the demand for compute expands. And I think that we're continuing to underestimate it. Do you worry about model training and model inference becoming, I don't know, 10x more compute? I think that generally speaking, what you're seeing is that if, let's say, you do become 10 times more efficient, I think that that just means that everybody is able to process 10 times more tokens, and there's still the same fixed amount of compute in the world at any given point in time.
What if AI gets 10x more compute-efficient?
Stephen Balaban [10:11] In the early days, it's funny, we used to talk a lot about this back in, let's say, 2017. Oh, well, maybe there's going to be some new type of model, let's say, that will look more like a random forest model, which the audience might—some members of the audience might know—you can kind of train a random forest model on a MacBook, right? And there was this concern that was kind of persistently raised around, like, well, okay, what happens if you have this sort of adjacent disruption on the model side of things?
The real bottleneck: land, power, and shell
Stephen Balaban [10:44] And so far, we haven't seen that. And again, everything that we're building towards is sort of based on these scaling laws, which is really about scaling up this architecture. So I don't really foresee a very likely outcome where we have this huge model disruption that would cause a decline in the demand for compute.
Matt Turck [10:53] Where's the main bottleneck these days that you're experiencing building Lambda? Is that GPU power? Electricity?
Stephen Balaban [11:09] So I always say that bottlenecks are always kind of local before they're global, in terms of one development might be bottlenecked on, let's say, generators or on UPS systems as a function of the sort of idiosyncrasies of the site. But broadly in the industry, the thing that is the main bottleneck is basically land, power, and shell, which is basically land that is entitled to have a certain amount of megawatt commitment from a utility, and then, of course, the data center and the mechanical, electrical, and plumbing equipment, the MEP equipment, that goes into that data center.
The backlash against data centers — and the misinformation
Stephen Balaban [11:38] And so that's the main bottleneck that we're seeing in the industry right now, I'd say, across the board.
Matt Turck [11:47] How real is the movement against data centers from the global community, and how do you think about how to respond to it?
Stephen Balaban [12:17] Well, certainly it's very popular in the news right now. I'd say that it's definitely very real. I mean, I think that, rightfully, communities that host any type of large capital project, whether it's a power plant or a solar farm or a data center or a distribution center, those communities want to have a seat at the table. I'd say, in general, though, I spend a lot of time reading through a lot of the comments from communities, and people want jobs, they want tax revenue.
Stephen Balaban [12:50] Any major capital development is going to bring a lot of tax revenue, and it's going to bring a lot of jobs. And it's going to bring investment into their community. And what they really are voicing, I think, is, one, having a seat at the table while this stuff is being developed. I think that's an important thing, just to have their voices heard and for the developers coming in to actually understand the community. The other thing to keep in mind is that there's a lot of misinformation out there.
Stephen Balaban [13:29] So, for example, every single modern deployment of, let's say, a Blackwell-class or a Rubin-class GPU, the GB200 NVL GPUs, these are oftentimes in a closed direct-to-chip liquid cooling system that's connected to a dry cooler, which means that there's almost zero evaporation. It's not using evaporative cooling. It's using a dry cooler system that does not consume a lot of water. And on top of that, most of these data center developments are bringing a ton of power to the grid. They're either standing up behind-the-meter power, they're standing up and bringing battery energy storage systems to the grid, and they're bringing all these sort of ancillary benefits that strengthen and fortify the grid and also eventually, in the long term, will maintain the costs that are being experienced by the community.
Stephen Balaban [14:23] And so I actually think that there's a very clear path toward maybe spreading more of the facts around what a data center brings, because there's just a lot of misinformation. You'll see people talking about how data centers consume a lot of water. Well, an evaporative cooling tower might evaporate a lot of water, but practically no new builds in the United States are using evaporative cooling for these closed-loop direct-to-chip liquid cooling systems.
Matt Turck [14:39] Do you think we do a terrible job as an industry explaining this to the broader world? Because those things keep coming back and they seem to be accelerating. But then when you have the discussion, from a technical standpoint, a lot of it is simply based on misinformation, as you just said.
Opening the hood: from photons to tokens
Stephen Balaban [15:00] I think that everybody's trying to get better at that kind of communication. And it just takes some clear thinking, writing down what are the benefits, writing down what are the costs, and presenting that clearly and plainly to a community so they can make a good decision about what kind of jobs and what kind of development they want in their communities.
Matt Turck [15:13] Let's open the hood for a minute. People talk about things like FLOPS and GPU hours and tokens and MFU. What is the best way to think about a compute unit?
Stephen Balaban [15:38] Yeah, it's interesting. You said a few different terms, and I always like to kind of break it down from a physics perspective into the SI terms. So, okay, on the left-hand side is all of the energy production, and then on my right-hand side is sort of tokens being consumed by somebody. And maybe you can even have the application layer on the far right of that that's using the token. So on the left-hand side, you've got either photons coming in per second or molecules of natural gas coming in per second.
Stephen Balaban [16:10] And then that, through a power plant or a solar farm, gets converted into joules per second, which is a measure of electrical power production. And then the joules per second—obviously, in engines there's a level of efficiency, and that's engine efficiency. It's interesting because the MFU percentage is kind of like an efficiency up on the higher end of that chain. The power plant or the solar plant then converts that into joules per second, which is watts, which is consumed by the entire data center.
Stephen Balaban [16:42] The data center itself needs to cool itself, and that's the PUE, and that's actually the efficiency metric that you can use to measure a data center on. And then you put the servers and all the different networking and storage gear in, and that's producing floating-point operations per second, or FLOPS. Okay. That is what gets consumed. The FLOPS capacity is what gets consumed by, let's say, a model builder when they're training a model or when they're inferencing a model.
Extracting more value from the same chip
Stephen Balaban [17:11] And that gets turned from FLOPS per second into tokens per second. Then, on top of those tokens per second, you might have some level of efficiency where the end customer is actually turning those tokens into real, actual intelligence. That's the entire pipeline, I would say, from end to end.
Matt Turck [17:24] Super helpful. If two companies have the same chip fundamentally, how do they extract more value from it? What needs to happen to maximize the usefulness of that chip?
Stephen Balaban [17:52] If you look at the cost structure of, let's say, one GPU-hour of time—we're talking about H100s—the largest part of that cost structure is the depreciation that is associated with that GPU hour. And basically, you can think of a utilization metric as being kind of a multiplicative factor on that. So, one over the utilization times the amount of per-hour depreciation expense associated with that. And so I think that the number one way that companies are sort of gaining a unique advantage is: how can I build a cloud product that is beloved by people that is going to drive high utilization?
Stephen Balaban [18:29] And in addition to that, the market, as we mentioned earlier, for on-demand compute—the retail pricing—is obviously much higher than the wholesale pricing. So retail is like on-demand: spin up a GPU, spin down a GPU, normal cloud service. Wholesale is sort of buying 10,000 GPUs for five years, for example. And so one of the things that we do at Lambda is really try to figure out, hey, how can we sort of get the most dollar utilization and percentage utilization out of the capital deployments that we do?
Stephen Balaban [19:08] And that's by making great cloud software that makes it easy for somebody to spin it up and down. So, for example, if you don't have that cloud software, you can't extract retail pricing, right? You cannot rent it out to somebody for an hour because you simply don't have the means to be able to do that. And actually, a lot of neoclouds are in that position where they don't even have the infrastructure to be able to run a real cloud service.
Frontier inference and distributed training, explained
Matt Turck [19:37] So you have GPUs, but a big part of how those data centers work is transforming GPUs into networks of GPUs. Do you want to explain at a high level how that works?
Stephen Balaban [20:17] The general idea is that you've got a large-scale, high-performance computing cluster of a bunch of, let's say, NVIDIA GB300 NVL72 racks. That's 72 GPUs all networked together via NVLink. And then there's a connection between the racks that's either InfiniBand or high-speed Ethernet. And that is essentially what's called a spine-leaf topology, which is basically a way to say, hey, this is completely non-blocking. Every port on every GPU can talk with every other GPU in the network. It's fully connected and it's able to provide maximum bandwidth between every individual GPU.
Stephen Balaban [21:03] And that cluster is useful for training large models. It's also useful for inference. So frontier inference, as we sometimes refer to it at Lambda, is basically very much a distributed inference problem where they actually will fragment or shard the model. There'll be some sort of sharding strategy for the model where it can be essentially run on multiple GPUs, and it uses that high-speed InfiniBand or Ethernet interconnect to do that communication.
Matt Turck [21:11] And so what is frontier inference? Is that inference for the most advanced reasoning models, like the more demanding jobs?
Stephen Balaban [21:33] It's not necessarily associated with reasoning models so much as just a very large frontier model that is kind of the domain of, let's say, three companies in the world or four companies in the world. When they're doing their inference, it's a very complicated thing that is fully utilizing all of the interconnection that's available.
Matt Turck [21:48] And what you described for frontier inference, is that conceptually the same thing as what happens for training? This concept of just distributing a task massively across a bunch of GPUs? What happens during a training run from a compute standpoint?
Stephen Balaban [22:12] Generally speaking, when you're doing a training run, you might think there might be some sort of split between the backward pass and the forward pass on the model. And the backward pass might be, let's say, two-thirds or more of the compute, and the forward pass, which is basically the same thing as inference, is the remainder. One of the realizations that I think has been made over the last bit of time is that the type of infrastructure that you'd want for doing a large-scale training run can be reused to do the inference of that model.
Stephen Balaban [22:56] What I mean by frontier inference and the fact that the inference is being done in a distributed way, you'll have, like, a mixture-of-experts model, and there'll be different, basically, sharding strategies for how you put those experts onto different servers and different GPUs. And the models can be very large. They may not fit on one single rack, or they may not fit on one single server. They might need to be distributed across different servers to even just do the forward inference pass.
Stephen Balaban [23:23] And so that's where distributed frontier inference comes into the picture. Because if you're doing a small model, let's say Llama that the users might be familiar with, or some of the quantized small models can fit on a single GPU. GPT-5 can't fit on a single GPU.
What actually drives compute cost
Matt Turck [23:45] And when we think about compute costs, what costs the most money? Is that model size? Is that memory bandwidth? Is that latency? Do context windows and those very large context windows change anything to the compute cost? What costs the most money?
Stephen Balaban [24:07] As I mentioned, the biggest component of the unit cost for a cloud service like this is the depreciation expense. And within that is basically some sort of bill of materials for the servers that are in the data center, which is by far and away the biggest portion of the cost. If you were to talk about the capital stack, let's say you can go back down to power generation: $2 to $3 million a megawatt, $2 to $3 billion a gigawatt for a power plant.
Stephen Balaban [24:55] The data center is between $10 and $15 billion a gigawatt for building the data center. And then the compute, the servers, can be anywhere from $35 to $45 billion a gigawatt. And within that, you can see the server portion is obviously by far and away the largest, and that's a big part of the depreciation expense. And then within that, obviously, you have the server and cluster bill of materials, which is primarily the GPUs. If you were to kind of break down NVIDIA's bill of materials, then you can kind of get better allocation toward where those costs are coming from.
Stephen Balaban [25:15] But certainly in the most recent period of time, memory expenses—memory has gone up a lot in price.
Lambda's chip stack and the NVIDIA relationship
Matt Turck [25:33] And there's very few vendors for HBM memory, but Samsung, Hynix. So you guys are a big NVIDIA shop. At a precise level, you mentioned some of the names, but which chips do you use mostly? What's your kind of chip stack?
A multi-silicon world? CUDA, CUDNN, and NVIDIA's real moat
Stephen Balaban [26:17] Yeah, so Lambda really loves NVIDIA's products. They're the only server provider, the only chip provider that is available in every single major cloud platform, which is a huge platform advantage. And we stuck with the NVIDIA ecosystem for all of the chips we've deployed. And we've got everything from V100s, A100s, H100s, H200s, B200s, GB200s, B300s, and VR200s coming soon. And so we use everything in the ecosystem.
Matt Turck [26:28] Do you think that today or in the near future, we're going to be in a multi-silicon kind of world? Is there room for different players beyond NVIDIA?
Stephen Balaban [26:52] Well, I think that we're already in a world where there's a huge amount of competition from massive, massive multi-trillion-dollar companies, and they're all trying to fight for the same thing, which is to be the best chip in the world for running and training neural networks, essentially. NVIDIA's built a great product that has gotten a lot of distribution and has a great platform of developers who love what they do. And you have to take into account not just the cost of the chip, right?
Stephen Balaban [27:22] The price of the chip is one aspect, but you have to take into account the entire software ecosystem and what's been developed. So one of the big things people talk about is, well, what's NVIDIA's moat? One of the big moats they've got is just the cuDNN stack. It's not just CUDA. CUDA, sure, that's like the water we all swim in, but cuDNN has got so many matrix multiplication routine optimizations baked into it.
Matt Turck [27:25] What is cuDNN for everyone to understand?
Stephen Balaban [27:35] Okay, so cuDNN is CUDA Deep Neural Network Library, and it's basically NVIDIA's—you can think of it like a highly tuned engine for matrix multiplication. And basically, if you were to just sort of naively implement the matrix multiplication algorithm, you would maybe get a certain level of floating-point operations per second, but they've gone and tuned every single aspect of it and come in and do Winograd filtering or a bunch of different algorithms that you would apply to speed up matrix multiplication.
Stephen Balaban [28:09] And cuDNN means that you don't have to go and do the optimization yourself. And so that's one aspect. The other one is NCCL, which is their networking optimization library, where it will sense the topology and the connected nature of your network, your InfiniBand or your Ethernet network, and it will suggest an optimized routine for doing all-reduce and broadcast, the different what are called MPI primitives, which are used for that sharding that we were talking about for both training and for inference.
Networking, storage, and the one-click cluster
Stephen Balaban [29:00] And so that's the kind of software stack that I think really is hard for a lot of the new entrants in the chip space to overcome. I think we're already, like I said, in a world where there are multiple options for silicon. The biggest labs in the world are using multiple different types of chips to do their inferencing and training on.
Matt Turck [29:09] What would be a plain-English definition? We talked about the chips, but the rest of the stack, the networking and the storage, just walk us through how it works.
Stephen Balaban [29:36] When you're running a cloud service, one of the things—you'll train your model or you'll upload your trained model, and you're ready to start doing large-scale inferencing. Well, you're gonna need a place to put your data, whether it's the data that you're using to train with, or whether it's the data that's coming in and streaming in from your end customers. And so having high-speed storage is a really important part of it. And so Lambda offers the AI-optimized file system service that is significantly faster than your standard, let's say, cloud file system, which is maybe more of a traditional NFS-type of thing.
Stephen Balaban [30:01] This is a highly optimized parallel file system that's designed for high-performance reads and writes, and mostly high-performance reads. That's kind of most of the workload.
Matt Turck [30:04] And that's something you built in-house completely?
Stephen Balaban [30:33] We have. I mean, it was in-house completely, right? You have to ask the question: what is the definition of in-house completely, right? We've never spun a PCB at this company. We have not authored—for example, we use KVM/QEMU for our virtualization. And so we have both commodity off-the-shelf hardware that has software installed on top of it for some of our storage. We have some storage partners that we work with as well.
Stephen Balaban [31:03] But generally speaking, everything that we do on the cloud, I would generally say, is something that we rolled ourselves with the help of the broader ecosystem. Because again, there's no such thing as rolling it yourself unless you're mining ultra-pure silicon from somewhere and then coming up with your own ASML. It's funny.
Matt Turck [31:11] Yeah, that's the highly optimized storage. What else? The networking part and what other pieces?
Stephen Balaban [31:39] Yeah. So I was talking about this one-click cluster product that we've got, and the way for everybody to think about this is, okay, well, look, you've got a bunch of GPUs. Let's say you've got a cluster of 10,000 GPUs. Well, I want to partition that cluster up. And so what it is, is it's a bunch of GPUs, some CPU servers as well, because you need to have an orchestration fleet as well. And then you've got some storage, and all of the CPU servers and the storage servers and the GPU servers are interconnected with the storage so they can quickly read and write from it.
Stephen Balaban [32:29] And that communication happens over what's called the in-band network. And then there's the compute fabric, which is where I was talking about all of the model weights and feature activations being shared throughout that compute fabric. And then there's an out-of-band monitoring network where you've got access to whether it's BMC or some of your DPUs. And when you are trying to create a subpartition of a 10,000-GPU cluster, you need to simultaneously partition the in-band, the out-of-band, and the compute fabric.
Matt Turck [32:29] Okay.
Stephen Balaban [33:02] That complex coordination between, we've got a bunch of bare-metal systems to, hey, we've got a virtualized system that has what's called RDMA, remote direct memory access, that allows them to read and write quickly, not just from the disks, but from each other's memory, the GPU's sort of HBM memory, and allow them to do that sort of direct memory access, allowing it to go directly from a GPU to another GPU without getting copied to the CPU, for example. Having that all work is an immense, immense software undertaking.
Stephen Balaban [33:36] And this is going back to the original question: what are people not getting about neoclouds? Well, first of all, the answer is that most neoclouds don't have this kind of technology. Most neoclouds have not made the high tens to hundreds of millions of dollars of software investment that you need to make to build a real cloud system that can partition a high-performance computing environment like this, and then to have it all work with the storage.
Stephen Balaban [33:52] Anyways, I guess that sort of summarizes the steps that you need, and you can think about all the different moving parts of a modern—how does an AI data center work? People talk about AI data center, but really you have to go down that one next level down, which is—because if you were to ask an AI data center landlord, a traditional one, what's going on inside of the data center, they'd be like, well, look, we're real estate people and we really outsource this to the GC, but the GC doesn't know.
Renting vs. owning, and full vertical integration
Stephen Balaban [34:46] Of course, anything that's going inside, it's their tenants who know. So this is what's actually happening inside of an AI data center, and then it serves the result. Also, going back to the community stuff, if people knew a lot more about, well, this AI data center is actually just serving the ChatGPT requests that I'm giving it, right? Sometimes they don't even realize that that's actually what an AI data center does.
Matt Turck [34:56] Yeah. So you mentioned tenants. Do you rent them? Do you also own some buildings? And where does that fit in the overall strategy?
Stephen Balaban [34:56] So initially, we started off as being primarily a renter, and we've actually started to get into the business of financing some of them, the construction of them ourselves, as well as we're going now into full vertical integration, where we are identifying land, coming to the table with a basis of design, which is basically all the engineering diagrams to construct the data center, financing and constructing that data center, putting the servers in, and then associating that with a long-term offtake agreement with one of the major compute consumers in the world and financing it all.
Stephen Balaban [35:51] So we're getting into full vertical integration at Lambda, and it's been great because we've been able to kind of, again, bring that engineering mindset to this problem, which was historically mostly run by people in real estate.
Matt Turck [35:58] In your own data centers, are you the sole tenant, or is part of the idea that you can also rent some to others?
How global is Lambda? Does location still matter?
Stephen Balaban [36:24] In a lot of our data centers, we are the sole tenant. In terms of the data centers that we're planning on constructing, we don't yet have any plans to lease that space to others. So we're not trying to get into the data center leasing business. Maybe that's something that you can imagine down the road. I wouldn't rule it out completely, but for now, we have to focus on providing Lambda with the compute that we need to service the market.
Matt Turck [36:27] How international are you, by the way?
Stephen Balaban [36:52] I'd say that we're very much focused on North America. And so we have data centers in Canada, the United States, and Mexico. We're very much, like I say, primarily focused on North America, but really within the United States, obviously. And we haven't had this desire internally to try to go and expand into Europe or too far into Asia. We've done some partnerships with some of our great investors, like SK Telecom, and we have a data center that we've operated in Korea, in Seoul.
Stephen Balaban [37:15] And so we have some experience with international, but right now we're just like, look, let's focus on the U.S. market. It's where the opportunity is.
Matt Turck [37:21] Do you need, for performance reasons, to be close to the customer the way you need to have regions in cloud?
Stephen Balaban [37:51] It's super interesting. I get this question a lot, and people are like, well, does latency matter? So I'll tell you what matters and what doesn't matter. You can look at your own utilization, whether it's ChatGPT or Claude or Groq or Gemini, and you can see, hey, a lot of the things that I'm doing, I kind of shoot it off, I come back later, and there's a research report for me. Maybe it's a long-running agent workflow. In those cases, latency doesn't matter at all.
Stephen Balaban [38:25] The only thing that matters is your cost per token. That's all that matters. And so that's been a really interesting change. I think the old-school, traditional legacy cloud business was so latency-focused because of some of the applications. But this new fleet of AI applications is far less latency-sensitive. So that's one. But there is the caveat, which is this: governance and data governance are becoming important things. And a lot of countries are wanting to have the AI compute that their citizens are using be run out of their own country so that they at least have their perception of control or whatever.
The financing stack: off-take agreements, SPVs, and credit
Stephen Balaban [38:44] And that is an element to it. But I'd say that from latency, there are no technical reasons.
Matt Turck [38:52] Let's talk about the financing stack. So presumably it's a combination of equity and debt. How does it all work?
Stephen Balaban [39:20] Yeah. So the way that it works is that you could really fragment it into these two parts, which is financing your on-demand cloud versus financing an offtake agreement, which is a longer-term commitment. And on the on-demand cloud, you're looking at Lambda's credit quality. On the offtake agreement, you're looking at the credit quality of the end customer who's paying the bill. And so what you do is you just take your offtake agreement, you take this chunk of GPUs that you're deploying, you take a lease on the property, and you kind of put it into a box, and you can go to the private credit markets and you can come up with an asset-based loan.
Stephen Balaban [40:14] There's a variety of different methodologies for financing it. Most of it is just some sort of special purpose vehicle that's designed to finance this particular deployment, with a very known and easy-to-underwrite— which is basically just a fancy way of saying the finance term for assessing the risks and the downsides of a particular credit investment. And there's a vibrant credit market for that. On the on-demand cloud side of things, it's not quite as mature as when there's, for example, an investment-grade offtake agreement.
Stephen Balaban [40:57] But it's becoming more and more mature. And in general, creditors and lenders are really starting to understand the value of an NVIDIA chip. Because you actually look at the chips that we deployed in 2023, H100s, we're now leasing those out at a higher rate now than we were originally in 2023. So these creditors are starting to look at these assets and say, wow, this is an asset that is very valuable and also easy for us to underwrite. And of course, while they are underwriting toward the actual cash flows that are coming out of that agreement, just as an asset class overall, people are realizing that this is a really great opportunity.
Why a 2023 GPU leases for more today
Stephen Balaban [41:16] And so creditors are starting to flock to these deals.
Matt Turck [41:33] You're renting an H100 at a higher rate because why? Because the demand for compute is so rabid that people will take any, or the technical depreciation of the product is slower than people thought. What drives that?
Stephen Balaban [41:56] Well, what's driving it? Certainly, it's the demand being high increases the price that you're able to get in the market. There's no question about that fundamental law. Again, going back to what people didn't understand about—there were people who were saying, "Oh, well, there's a five-year lifetime or three-year lifetime." I even heard some people say three-year lifetime for these GPUs. This is completely false. We have GPUs that we've commissioned, and we're one of the earliest neoclouds.
Stephen Balaban [42:23] In fact, we're probably the only neocloud that actually has GPUs in our fleet that are fully depreciated from an accounting perspective, right? Most people are adopting around a six-year accounting depreciation schedule. But that's not the usable life. The usable life is longer than the accounting depreciation schedule. And what really matters is the economic usable life. And so what we're starting to see is that the people who are the naysayers—"Oh, this is going to be—you’re going to throw these GPUs out in five years"—are completely wrong.
A futures market for compute?
Stephen Balaban [42:36] They're completely wrong, and they've been wrong the entire time.
Matt Turck [42:49] Do you think there is going to be, or do you already see happening, some kind of financial market for compute units with trading and derivatives? Is that happening?
Stephen Balaban [43:22] I'm starting to see some people start to examine what a maybe vibrant spot market—first, you need to have a spot market for something before then you can establish a derivative like a future or other more exotic things. I'm starting to see that, but fundamentally, I think that the asset class is just starting to mature, and creditors are starting to become very comfortable with investing in the credit side of buying NVIDIA GPUs and deploying them into data centers. And we don't need to get too fancy with it.
Origin story: facial recognition, Perceptio, and Apple
Stephen Balaban [43:54] That's kind of my opinion, is that I think that market is starting to mature, that maybe an eventuality is having more complex securities that surround GPUs. But I think for right now, people are starting to realize that it's a great credit investment. And that's what's changed, I'd say, over the last year, is that people have started to really treat it like a more mature asset class.
Matt Turck [44:06] Maybe quickly just go back to the very origin, because I think you've been effectively in the AI world the whole time, but are coming from a very different angle with multiple pivots. What did you start with, and when?
Stephen Balaban [44:38] Well, with the complexity of the business, you can now see the complexity, the capital intensity, just the sort of not fitting into a box. And you can see why we've oftentimes not had a lot of traditional venture investors in Lambda. And all of our investors have done exceptionally well, but they've kind of come from, more often than not, outside of traditional, let's say, mainline Silicon Valley VCs. And so, just going back to the origin story, I started Lambda in 2012, and we were a facial recognition software company.
Stephen Balaban [45:19] So, I was training convolutional neural networks to do face and image recognition. And we eventually hosted that on an API. I was training those ConvNets on a four-NVIDIA-GTX-580 workstation that I had bought from a friend who had built it, actually. And this was really pretty avant-garde stuff at the time. Most people didn't really believe in the field of deep learning at the time.
Matt Turck [45:24] And that was inspired by the ImageNet 2012 moment, or that was even before that?
Stephen Balaban [45:48] The ImageNet moment, I pulled the cuda-convnet repo off of Google Code. That's how old Lambda is, is that Google Code was still around, and I pulled the cuda-convnet codebase and was playing around with it. I got very lucky that the AlexNet paper had been published the same year that Lambda was founded. It's not a coincidence at all. It's not a coincidence at all. We launched this face recognition API, got a couple thousand users, but it wasn't really generating a ton of cash.
Stephen Balaban [46:21] And sort of as part of that complex story of startups, in parallel, I found these guys who had just graduated from their PhD programs, Zak and Nico, and they had said, "Hey, we're going to start a company." And I said, "Hey, let me help you guys out. I'm going to work with you for a year. I'm going to learn a little bit more about neural networks." And I helped them out on this company, helped them get a company called Perceptio started.
Stephen Balaban [46:48] And I was the first employee there while I was running Lambda. And we were running these ConvNets locally on the iPhone. And again, this is 2013. So we were using the GPUImage library and just straight OpenGL ES shaders, like the shaders that are used for rendering. We were using those to run the ConvNets on the iPhone. And eventually, I kind of left to go continue to work on Lambda full-time, and probably about a year or so later, they got acquired by Apple.
The Lambda hat and Dream Scope
Stephen Balaban [47:18] And so, if you know the feature on your iPhone where you swipe up on an image and you can recognize faces and search through your library, that's maybe some of the stuff that eventually got integrated into iOS through that acquisition. And then Lambda, we continued on. We had a variety of different products, everything from Lambda Hat, which was a baseball cap that took a picture every 10 seconds with a camera embedded in the tip of the brim for gathering datasets for image and face recognition.
Matt Turck [47:38] Which is fascinating because fast-forward to today, and that's a whole segment, right? Capturing everyday life to train the AI.
Stephen Balaban [48:05] It goes to show you have to—one, it's important to be able to see the future. It's also important to get your timing right as well, right? And now it all worked out, right? Despite maybe that Lambda Hat product not being great, but it taught me a lot about how to build hardware. I lived in Shenzhen for a little bit, working on the PCB and spinning the PCB and designing the actual hardware product. And it taught me how to make consumer electronics, and that was actually a huge, huge skill because it totally opened my mind to
Stephen Balaban [48:42] new ways of doing business that aren't just making apps, right? And eventually we had this product called DreamScope, which became really popular in 2015 and '16. And it was basically using the Google DeepDream methodology of using a ConvNet to generate images. It's like an early version of Midjourney or whatever. And DeepDream and the Léon Gatys style transfer algorithm allowed you to turn a photo into a painting, basically. And we got like a million users on that, processed tens of millions of images, maybe 15 million images or something like this.
The $60K bet that became a cloud business
Stephen Balaban [49:18] And that caused us to have a huge AWS bill. It was like $40,000 a month or something. And so, to replace that, we ended up building a little cluster out of workstations. And then there was a $60,000 CapEx that we were terrified to make, by the way. We were so scared that doing this CapEx was going to put us out of business. We made it out of workstations because we thought, oh, well, worst-case scenario, we can just sell them. And so, lo and behold, we did end up turning it online, and it brought the bill down to zero.
Stephen Balaban [49:52] So it paid itself back in a month and a half. And we thought, wow, this is like, we're saving more money than we're making. Maybe we should be in the business of providing compute to other AI researchers. And thus, we started selling workstations and servers and started developing a cloud platform. Maybe did $3 million of revenue in 2017, that first year selling workstations, then $10 million in 2018, then $30 million in 2019. We grew the hardware business over the next couple of years to probably about a $200 million run rate.
Stephen Balaban [50:27] And then the cloud business, we really started in 2019, and we started development before then, but we started really marketing it. And Lambda was slow to grow, to be honest, because not a lot of people in 2018 and '19 and 2020 wanted a bunch of AI compute. There was a pretty niche market for it. But eventually, our cloud business continued to grow, and now it's at a little bit under a billion-dollar revenue run rate. We've fully exited the hardware business.
Stephen Balaban [50:34] And so, yeah, Lambda's got an absolutely wild founding story, to summarize.
Matt Turck [50:42] Are some of the people that were there at the beginning still around? I think you started the company with your brother, is that right? And your brother is still at the company?
Stephen Balaban [51:14] Yeah. And so, in terms of the early people, of the four people who were making Dreamscape—me, Michael Balaban, my co-founder and fraternal twin brother, Chuan Li, who's our chief scientific officer, and then Steve Clarkson, who's an engineering leader at the company and has a bunch of folks reporting into him now—they're all still at the company. The next hire, one of those gentlemen named Mitesh Agarwal, who was one of the next hires in that team, he was with the company for maybe eight years or something like this.
Holding the team together through the hard times
Stephen Balaban [52:00] Yeah, something like eight years, and then he eventually left and joined another former Lambda team member, Thomas Summers, to start Positron, which is an accelerator company, and they're now valued at over $1 billion. And so not only has the original team stuck around, but we've already started to kind of see what a Lambda alumni network looks like in the world.
Matt Turck [52:05] How did you keep the band together during the difficult times?
Stephen Balaban [52:33] Just when you're running a startup company that's this capital-intensive, working-capital-intensive as well, you get a lot of shocks to the system as you're growing. And then COVID—I mean, in April of COVID, software companies were feeling great because they could ship software and there was so much more demand. And hardware companies, the docks were closed. You couldn't ship revenue in April and March. I remember all these things really distinctly.
Stephen Balaban [53:01] I think I remember just getting in front of the team like, hey, look, it's really tough right now. And there's certainly a feeling that we're not sure if we're going to make it through this. The only thing to do is just to suck it up and enjoy the pain, run through it, and come up with the solutions to the problems that you're presented with, all in the service of delighting customers. Because fundamentally, the big thing is just aligning people toward the only reason we're all here: to build something that people want and they love so much that they tell their friends about it and they give you money.
Stephen Balaban [53:43] And then everything else, it just follows from that customer experience of delighting customers with what you do. When we do onboarding, for example, I used to do this thing called Lambda 101, and we would show a picture of a Linux penguin, and he was on a Lambda workstation and he was reading the GPT-2 paper, and training had a loss curve, which is what you see and look at if you're doing machine learning research. I was like, just put yourselves in the shoes of the penguin who's using this workstation or cloud service to train a neural network, and just think about what's gonna delight them.
Stephen Balaban [54:19] Whether it's people on our shipping team who said, hey, let's put some T-shirts inside of the boxes, and so every workstation came with a Lambda T-shirt. Or members of the data center operations team said, hey, we should do a white rack because that'll kind of set us apart and make everything look good, and we'll be really proud to showcase that. And those are the types of things that, as you kind of imbue your company with the delight-the-customer-first mentality, I think help you get through the hard times.
Bringing on a new CEO; Stephen as CTO
Matt Turck [54:46] Recent evolution in that journey is that you just brought on a new CEO, and my fellow French countryman, Michel Combes, to run the business. Walk us through the thinking and what led you to make the decision and how that equips the company for the next chapter.
Stephen Balaban [55:17] It's a huge honor as a founder to get to the point where the company can afford to bring on amazing talent like Michel in that seat, right? Just because, if you think about it, I'd say most companies—it's not uncommon for somebody to say, hey, look, a lot of people, sometimes maybe there's a component of ego involved where they have to be the founder CEO. I've never really personally had that. I care about the technology, as you can tell.
Stephen Balaban [55:40] I care about building a great generational company. And I think there's so many different seats to do that from. And so, getting to the point of maturity where we could afford to bring on a CEO like Michel, who has experience, obviously previously SoftBank International CEO, Sprint CEO.
Matt Turck [55:41] Alcatel.
Stephen Balaban [56:11] Alcatel. He's on the board of some really amazing companies, including McLaren, which is kind of a fun one. I always did the sort of fundraising and capital formation and day-to-day business management as a necessity and not because that's what I really love doing, for example, right? And I think there's plenty of founder CEOs who absolutely love every aspect of their CEO job. I think that privately—and it'll be very hard for you to get this out of any founder CEO oftentimes—but secretly, when I talk with founder CEOs, I'm always like, yeah, so how much do you hate this?
Matt Turck [56:33] I find it shocking that people don't find speaking to VCs all day exciting, but I will take your word for it.
Stephen Balaban [56:55] It's been, like, an amazing experience for me that I get to be able to form a team around the company and to just see everybody flourishing in the things that they love to focus on. So, for example, now that I'm the CTO, one of the main things I'm focused on is what does rapid data center deployment look like at the company? And kind of working to say, like, hey, I want Lambda to be this sort of vertically integrated, high-velocity powerhouse, so when you look at the world, you say, all right, there are three companies in the world that can do high-velocity deployments: SpaceX, xAI, and Lambda, where we're just extremely focused on how do you cut every little piece out of the process to stand up compute faster.
Matching xAI on high-velocity deployment
Stephen Balaban [57:39] And that's something I've just been diving into and really enjoying with my new time. What was xAI's record when they launched Colossus 2? I think it was like 200-and-something days.
Matt Turck [57:44] Yes. And you think that can be matched or exceeded at a repeatable pace?
Stephen Balaban [57:47] I think it can be matched or beat, yeah.
Matt Turck [57:49] And that's process, mostly?
Stephen Balaban [58:09] I think it's everything from the site selection process, the set of constraints that you use in a site selection process, the MEP pipeline, the way that you construct the data center. How do you make it so that the end customer will consume that compute? And how do you cut a lot of stuff out of the process? Because oftentimes, the people who've been designing these data centers have really kind of been real estate people, as I've mentioned, who've been kind of grabbed by the scruff of their neck by a hyperscaler.
Stephen Balaban [58:51] And they're like, "Go and build this design here. Go get a GC, run off." And they don't know anything about what goes inside of it. And the hyperscalers, on the other hand, have been really building towards traditional cloud services. I mean, if you look at a modern region in any of the clouds, they have hundreds of services. I mean, everything from satellite base stations to tape storage to spinning disk to face-recognition APIs. I mean, these are all the services, and each of those services requires a different SKU and has different parameters.
Stephen Balaban [59:23] Parameters about what you're kind of servicing. And in fact, you might have somebody who's trying to run an ATM backend on one of these things. That's a pretty different design space and design constraint than an AI data center that could maybe have lower availability and uptime, right? And so that's kind of where I think Lambda is able to build a lot of really unique value.
"AI won't write software — it will become the software"
Matt Turck [59:37] And through this, kind of, you had a quote where you said that AI won't write software; it will become the software. What do you mean by that?
Stephen Balaban [1:00:08] So that's in my sort of idea around what I call neural software and/or a neural computer, neural operating systems. And the best way to kind of get this experience is to go to your ChatGPT or your Claude and say, "Hey, just render for me an ASCII art desktop interface. So you're working purely in the domain of text, and I want you to just pretend to be an operating system for me. I'm going to say, click on this, open up this, and I want you to just behave like a computer."
Stephen Balaban [1:00:18] So give it that prompt.
Matt Turck [1:00:18] Okay.
Stephen Balaban [1:00:45] And what you're gonna see, I think, is that you're gonna see that sort of future of the large language model becoming the software and not generating the software. And this results in an extremely sort of squishy and flexible way of interfacing with a computer where it's not possible to have a bug, only a misunderstanding about the prompt and what you've asked for. And I think that for a lot of the pieces of software on your computer, you might see that taking over, where you can get the glimpse of the future with this ASCII art, and then eventually it'll also have a multimodal network that's generating every pixel on your screen as well as every audio waveform that comes out of your speakers.
Neural software vs. vibe coding
Stephen Balaban [1:01:31] The advantage to this is that you can really sort of dream up software, only the part that is being experienced by you and is actually implemented, if that makes sense. It could have whatever feature it is if you ask it. And that's a really powerful way to interact with a computer, I think.
Matt Turck [1:01:37] So it's not like you give simple instructions to the LLM and suddenly the LLM is the software.
Stephen Balaban [1:02:06] I guess, make the analogy: vibe coding takes in a prompt and then outputs human-readable, writable, compilable code that runs on normal human software programming language substrate. It outputs C code, which gets put through a compiler. It outputs Python code, which gets put through a Python interpreter. That software is static. Once it's been generated, it can't change. You can vibe code it again and maybe vibe code on the fly. There's a couple different stages of the gradient between traditional human-written software, and then you go to maybe vibe-coded software, then you go to just-in-time vibe-coded software, where it's a live creation of the software application.
Matt Turck [1:02:23] But still software.
Stephen Balaban [1:02:53] But still software. But then you go to the next step, which is just you're interacting with the LLM and it is emulating kind of how software might behave. And that's the difference between vibe coding and a neural operating system or neural software. Neural software, there is no code that's running. It's just modifications of the feature activation space and the context in the mind of the neural network.
Matt Turck [1:02:57] How far do you think we are from that? Is that something—
Stephen Balaban [1:03:01] I mean, we have prototypes of it today.
Matt Turck [1:03:04] When you say you, is that Lambda or is that others?
Stephen Balaban [1:03:42] Yeah, Lambda's developed a prototype. There are multiple other companies that have developed prototypes of this. There's academic research that has outlined what this might look like. And how far are we from mass adoption? I would say that, generally speaking, when I'm early on something, I tend to be about a decade to a decade and a half early. So I would say that between a decade and 15 years, we will see mass adoption beginning or otherwise happening for neural software. I mean, you already have it.
Stephen Balaban [1:04:01] So here's another example, by the way. You can think of a Tesla self-driving car, or any type of end-to-end neural network and then large model that's doing autonomy, as a form of neural software.
Matt Turck [1:04:02] Right.
Do agents change the compute layer?
Stephen Balaban [1:04:26] And people understand that aspect, right? Which is, it's seeing video, it's making decisions about what to output. Now, the user experience is the driving experience. That said, that is an example of neural software, I would argue. And so we already see that today. Now, the question is, when are everyone's computers going to adopt that? I'd say a decade.
Matt Turck [1:04:32] Do agents change anything from your perspective as a compute provider? And if so, in what way?
Stephen Balaban [1:05:03] To understand what needs to change on the compute layer, we need to understand what's changing with the user. When you're doing vibe coding with agents, one of the things you'll notice is that your wall-clock time in the world is mostly spent on running tests, gathering data, searching through the codebase. A lot of the time is spent not just inferencing a neural network, but it's actually spent doing other things. And it's actually very much similar to how software engineers spend some of their time.
Stephen Balaban [1:05:44] You know the old XKCD cartoon of compiling, where they're sword fighting on office chairs and someone says, "What are you guys doing?" "Well, compiling." So now there's a bunch of time spent compiling, there's a bunch of time spent running tests, because part of the way that the agent 24/7 loops really work well is when you are constantly banging against a nice suite of automated tests to make sure that the code you're writing is good. And so, well, what does that mean?
Self-assembling software inside Lambda
Stephen Balaban [1:06:14] It means that every single cloud service needs to start doing a lot more traditional CPU workloads. They need to focus on a great environment, a secure environment to host your Claude Code instance on. And then you need to think about security from the perspective of how this massive influx of new applications are going to be secured.
Matt Turck [1:06:17] How do you use AI agents internally?
Stephen Balaban [1:06:38] A lot of the engineers at Lambda are already doing a fully agent-driven workflow. If you just go to Claude Code and say, "Hey, use advanced workflows or spin up agents," you can do that. So that's, like, step one. I've demoed internally, and some folks have adopted what I kind of call self-assembling software. And so self-assembling software is this idea where you kind of tie into a 24/7-running agent fleet to meet product requirements and constant user feedback that's coming off of the system.
Stephen Balaban [1:07:03] So you have a very clear and tight loop to go from submitting, "Hey, this is a bug," or, "This is a feature request," and there's a fleet of agents who are implementing that live for you. And that sort of cycle I call self-assembling software because you kind of say, "Hey, this is what the software is for," but most of the development for it is going to happen after the software is launched and the users start to interact with it and customize it for themselves collectively.
Stephen Balaban [1:07:47] And I think that that is kind of maybe the future paradigm of where a lot of the agent-driven development is going to go towards. The other side of that, eventually, once the models get smarter—I think that they're not quite there yet—is tying that back into, "Hey, I need help." And I'm not talking about the human; I'm saying the agent's going, "Hey, I need a human to help me. Like, I need you to plug in 1,000 GPUs for me, or I need you to give me an API key to a particular service."
Gigawatt-scale AI factories
Stephen Balaban [1:08:18] I need you to go sign up for something for me. Can you please go negotiate this? And I think that that's actually how you're going to start to see it happen, which is product-user feedback gets implemented by the agents. The agents then also ask the people at the company to go and do things, all in the service of delighting customers and making money.
Matt Turck [1:08:32] You've talked about gigawatt-scale factories. Is that what you were describing earlier around setting up, getting super good at creating data centers very quickly, but also making them bigger? What is that concept?
One person, one GPU
Stephen Balaban [1:08:57] It's an AI factory, which is basically land, data center, servers inside that is generating tokens. And a gigawatt scale means that it's consuming 1,000 megawatts or 1 billion watts, which is a lot of power. Maybe you can think of it for context: New York City is something like 5 gigawatts.
Matt Turck [1:09:05] You also talked about one person, one GPU. Is that your vision for the future? Unpack that for us.
Stephen Balaban [1:09:34] So before people really believed in the AI thesis, when I was pitching our Series B and C, I would talk a lot about the similarities between, let's say, the computer industry and the AI industry. I really felt like AI was forming a set of generational companies, and there was going to be a set of generational companies that got minted with the changes that were coming with AI. And this is like in 2020, 2021. And if you read about the history of Apple, for example, in the early days, the motto and the credo at Apple was, "One person, one computer. One person, one computer."
Stephen Balaban [1:10:14] And there's a sense of humility that's embedded in this one person, one GPU, which is the one person, one computer. You think about how visionary Steve Jobs was. Apple was founded in 1976. The Macintosh came out in 1984, eight years or so after founding. Is that one person, one computer yet? No, not even close. So, 1984 to 1994: is it one person, one computer?
Stephen Balaban [1:10:46] Well, we're just starting to have the internet boom, so we're not quite there yet. 2004, we finally have broadband internet access. And maybe for the first time in the United States, there's not quite one person, one computer, but there's certainly like one person, one family, one computer, or something like this. It's getting close to it. You don't have until 2014. So '74, '84, '94, 2004, 2014, 40 years after one person, one computer.
Stephen Balaban [1:11:05] Do you have probably truly one person, one computer? And you get actually beyond one person, one computer because people have laptops and cell phones, and I would consider a cell phone a computer. And then finally, you didn't even have e-commerce penetration until 2024, 50 years after the founding or so of Apple Computer, when e-commerce starts to actually penetrate because of COVID. I think that the reason I really wanted to choose that one person, one GPU is because, one, I believe that in the future everybody in the United States will need the computational power of one GPU or more to just do their daily work, enjoy life, whether it's getting access to—whether it's getting entertained, whether it's being productive, whether it's being creative.
Hot takes: overrated and underrated in AI
Stephen Balaban [1:12:04] And I also recognize that it took Steve Jobs and Apple, one of the best companies in the history of capitalism, half a century to accomplish their goal. And so I think this is not just like an overnight, let's quickly get to one person, one GPU. So that's what that means to me.
Matt Turck [1:12:08] To close, are you ready for a couple of quick hot takes?
Stephen Balaban [1:12:09] Sure.
Matt Turck [1:12:13] What is one idea in AI that is overhyped?
Stephen Balaban [1:12:38] I think a lot of the sort of agentic workflows for things that are not software engineering tend to be overhyped. And I'll tell you that the reason for that is because one of the ways that you get an agentic workflow working really well is that it needs to have very concrete feedback mechanisms, which are done brilliantly through automated testing. It's not at all done brilliantly for going to buy a site. There's no gradient to give a model to go and iterate over a long period of time on.
Stephen Balaban [1:13:13] So I think agentic workflows for things that are readily verifiable. Now, I wouldn't say as far as everything that's not software engineering, because there's plenty of readily verifiable fields: CAD, computer-aided manufacturing, finite element analysis, computational fluid dynamics. There's a bunch of fields where you can really do a great agentic workflow and simulate it and then go and iterate. It's not the case for, hey Claude, make me a billion dollars, make no mistakes.
Matt Turck [1:13:17] Sadly, or maybe not. Okay, fascinating.
Stephen Balaban [1:13:24] It wouldn't be inflation; it would actually be just value creation in the economy, deflationary even.
Matt Turck [1:13:27] What is one idea in AI that is underrated?
Stephen Balaban [1:13:48] Yeah, I really think that the neural OS thing and also some of the aspects of self-assembling software. I still think people—the funny thing is, I'll give the same answer: agentic workflows for software development. I think that most people don't understand. They literally don't understand because they've never tried it. They've never gone to Claude Code. Go to Claude Code, say maximum effort, use the latest model, and then go and build whatever you wanted to build and say, spin up 10 agents to go and do it.
Stephen Balaban [1:14:00] I think a lot of people still haven't done it yet.
Matt Turck [1:14:04] Well, Stephen, it's been wonderful. Thank you so much for spending time with us.
Stephen Balaban [1:14:06] Matt, thank you so much for having me. Appreciate it.
Matt Turck [1:14:28] Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.