AI, Data and Blockchain: a VC perspective | Tomasz Tunguz, Founder of Theory Ventures

The MAD Podcast with Matt Turck · with Tomasz Tunguz, Founder, Theory Ventures

Tomasz Tunguz is the Founder at Theory Ventures. We cover why $80 billion in AI investment could account for nearly half of venture funding, why Web3 databases can justify higher costs for cryptographic guarantees around sensitive data, and why routing routine queries to specialized small models can be more practical than relying on one general-purpose LLM.

Watch on YouTube

Chapters

  1. 2:46 — Tomasz has continued to invest in blockchain through the crypto winter. Why?
  2. 6:59 — Security and privacy as the main blockchain's use case.
  3. 9:18 — Blockchain and AI: how do they work together?
  4. 11:02 — Why does Theory Ventures not invest in AI hardware?
  5. 12:28 — Why do big companies invest in cloud infrastructure?
  6. 15:35 — An investor view on the foundation models.
  7. 18:36 — Is Gen AI going to replace traditional AI?
  8. 20:57 — Does the Theory Ventures invest in AI tooling companies?
  9. 22:53 — Is investing in Cloud companies better than investing in AI-powered applications?
  10. 26:40 — Copilot AI vs full-execution AI.
  11. 28:38 — A case for specialized LLMs.
  12. 29:54 — Gross margins in Gen AI: is it profitable?
  13. 32:34 — Modern Data Stack: is it still a thing to invest in?
  14. 37:02 — Microsoft Fabric and its impact on the market.
  15. 38:50 — Tomasz's thought on Motherduck and DuckDB.
  16. 40:37 — Where do BI tools fit in the Modern Data Stack?
  17. 44:32 — Why has the democratization of BI never happened?
  18. 45:52 — How do acquisitions happen? Can you engineer them?
  19. 49:02 — Key ingredients to build data infrastructure business.
  20. 50:40 — Tomasz is a founder now! How does it feel?
  21. 53:15 — Talking numbers: Theory Ventures' financial model.

Transcript

Tomasz has continued to invest in blockchain through the crypto winter. Why?

Matt Turck [0:51] Hey, Tomasz, welcome.

Tomasz Tunguz [0:53] Thanks, Matt. Pleasure to be here.

Matt Turck [1:05] So you and I are very much interested in the same topics, and weirdly, this is actually the first time we meet in person. So I'm excited. I'm particularly excited for the conversation.

Tomasz Tunguz [1:11] Yeah, it's bananas that we haven't met before. It seems like we've been on parallel coasts and parallel tracks. Yes.

Matt Turck [1:40] And finally, so on this podcast, we mostly have operators, founders, people building companies, but occasionally we have conversations among VCs, which I enjoy very much as well. So for anybody listening, this is kind of like two VCs shooting the breeze, talking about data and AI and gossiping about venture. And that's actually very much what VCs do when they grab the proverbial coffee every now and then. Except I think this conversation is going to be more actually in the weeds of data and AI, given we both do a fair amount of work in the space.

Tomasz Tunguz [1:51] And there's a lot happening there.

Matt Turck [1:55] Yeah, absolutely. Yeah, I hear AI is hot or something.

Tomasz Tunguz [2:16] I was running the numbers. There's a blog post coming out today. I think there'll be something like $80 billion invested in 2024 in AI. Wow. And if you think about venture capital, it hit its peak, I think, in '21 at $300 billion, and then it fell to $175 billion. So if we're anywhere close to the $175 billion number, you're literally talking about half of venture dollars going into AI. Part of it is that every company is now an AI company.

Tomasz Tunguz [2:21] As a keyword, it's in everybody's description.

Matt Turck [2:27] But I sort of—the question: where do the dollars—what other type of companies do the dollars go into?

Tomasz Tunguz [2:46] Yeah, but it'd be crazy if it were half, close to half of all venture dollars pursuing one category. I mean, it tells you one of two things. I think the market's really large and hugely disruptive, and the value creation will be enormous. But it also tells you that there's probably some pricing arbitrage in other categories where people aren't paying attention.

Matt Turck [3:03] Before we get into data and AI, that's actually partly what you do as well, right? I mean, not that blockchain and crypto is necessarily cold, but it's certainly an area where the heat has moved away from. But you've continued to invest through the crypto winter?

Tomasz Tunguz [3:25] Yeah. So I think, 18 months ago, if you had a great person leaving Facebook or Google, they were going to Web3, right? And this is all before the FTX disaster and the Fed raising rates and the economic environment completely changing. And then I would say when we started the firm last year, you could really name your price. You could really approach most crypto businesses and say, "We'd like to invest 10," and the post-money didn't really matter.

Tomasz Tunguz [3:58] And that was because a lot of builders left. Electric Capital puts out a report. They had 25,000 developers. I think that number's fallen to 15,000 in all of Web3. There are 27 million software engineers in the world. So we're really talking about a fraction of a fraction of a fraction. Yeah. And then before the Bitcoin ETF, it was really quiet, and it's super cyclical, just the way that the startup ecosystem is cyclical. I think crypto is that, but with greater amplitude changes.

Tomasz Tunguz [4:23] Now it's come back. I mean, I think hot seeds are 150. It looks an awful lot like AI. And then many of the token launches will raise equity rounds before they go public, and those will be in the several hundred million to billions again. So it's come back. It's come back in a really meaningful way. The number of players is much smaller. And I think one data point is that there are only three major crypto publications that write.

Tomasz Tunguz [4:50] Most of the journalists have actually left. They've needed to go to other publications because there's not enough money or there aren't enough stories to cover. So now there are three crypto publications in totality that matter for a fundraising announcement or a new product announcement. It's really small, but we're keen on it. So let's take Ethereum as an example.

Matt Turck [4:51] Okay.

Tomasz Tunguz [4:52] Six Snowflakes.

Matt Turck [4:55] The total market cap of Ethereum.

Tomasz Tunguz [5:16] Six times Snowflake's market cap in Q1. If you were to look at Ethereum as a business—and we can have this debate about whether it is or not—but if you were to look at it as a business, it produced roughly $400 million in free cash flow or net income, which makes it the most profitable software company on the planet: 44% net income margin. No one else is anywhere close. If you were to look at that total, that sum, that $400 million as net income, it would be the sixth-largest producer of profits of any publicly traded software company in the world.

Tomasz Tunguz [5:54] You have Microsoft and Zoom and others on top. And so the way that I look at it—and this is not a broadly held view—but the way I look at it is, that's a database company. And yes, there's programming associated with it, but Oracle databases have stored procedures. There's programming that exists. Snowflake has user-defined functions, Snowflake Snowpark. And so I look at these Web3 databases as exactly that, and they can be phenomenally interesting businesses, right? Ethereum being as valuable as it is, producing the amount of cash.

Tomasz Tunguz [6:21] The other way we look at it is that they're following a price-performance curve over time. If I were to write a transaction to an Amazon database like RDS, that would cost me a certain amount of money. Three years ago, if I were to write the same data to Ethereum, it would cost me a million times more to write that row. Today it costs about 1,000 times more. And so there's this asymptotic cost curve that will come down. And at some point, our belief is that most major software companies will actually have a Web3 database as part of their stack, particularly for very sensitive information where they need cryptographic guarantees.

Tomasz Tunguz [6:56] They either need the user to custody the information, or they need to show the German regulator that the data is stored on German servers, or that the users custody it. And so particularly if you're a software company with 80% gross margins, you should be willing to pay a premium for those kinds of guarantees. And you'll basically, we think, be offsetting the cost of compliance with a bit more expensive database.

Security and privacy as the main blockchain's use case.

Matt Turck [7:08] All right, interesting. So security and privacy of data is the use case internally and also to share amongst companies, or no?

Tomasz Tunguz [7:32] There will definitely be a data-sharing use case. I think we're starting to see that happen more at the developer tools layer, not so much at the Snowflake-Salesforce integration layer. That really hasn't happened yet. But we do think that security and data custody would be one of the big drivers. Because I think by the end of 2025, something like 35 states in the US will have their own privacy regulation. And so California, just the way we do with cars, we have different laws than everybody else.

Tomasz Tunguz [8:02] But then Dubai has different laws, and Germany has different laws, and South Africa will have different laws. So if you're a big company and you have multinational customers, the cost of compliance will become significant. And then there's also the honeypot effect, right? Okta is storing all of these keys, and so hackers are really interested in breaking that open for obvious reasons. But imagine if all that information were stored in a decentralized way.

Tomasz Tunguz [8:25] Fine, you break into one key, you have access to that one key. It's actually much better to build it in a decentralized way for those reasons. So, wiring information. The example that we give, that we talk about a lot, is: imagine if you were to build Salesforce on a Web3 stack today. So let's say I'm a vendor, I'm selling Figma to FirstMark, right? I'm an account executive, and I call you and say, "Great, Matt, thanks for signing the $100K deal."

Tomasz Tunguz [8:56] We need your banking information for whatever reason. Well, then you call that over the phone, and I put that in my Salesforce, and that information is now stored in all of your vendors' Salesforce instances. One, it's duplicated, and if you move, like you just have, you need to update your address, which doesn't make any sense. And then the second is, let's say you change your wiring information, you have to go and communicate all that. There's risk of wire fraud.

Blockchain and AI: how do they work together?

Tomasz Tunguz [9:18] But the second is, now there are 10 or 15 or 20 or 100 different places where somebody could access that information. Imagine instead if you had a wallet, or FirstMark had a wallet, that stored all its payment and address information and just gave access temporarily to a vendor for particular reasons and revoked it instantly. Much more secure, much safer, easier to manage.

Matt Turck [9:26] Are you at all interested in the frothiest of all frothy areas, which is the intersection of blockchain and AI?

Tomasz Tunguz [9:53] So we have looked a little bit there. I think the majority of the companies are GPU farms, I mean, distributed data centers for GPUs. Render is in that category. And I think they were just used on the new iPad Pro demo yesterday. One of the games was using the Render engine. I think the big question for those companies, just like the distributed GPU, is latency. Latency really matters, particularly inference time. And so it seems to us that most of those clouds were probably used for training.

Tomasz Tunguz [10:22] And then the question is, what is the relative arbitrage opportunity between a decentralized GPU network versus a centralized GPU network? Data movement, ingress, and egress are significant. And then one thing that we learned yesterday was that when you rent a GPU—not one GPU, hundreds of GPUs—in order to train or infer, it's really expensive. And so you want to move the data as quickly as possible in so that those GPUs are saturated with the data and then they can infer really fast.

Tomasz Tunguz [10:43] If you're spending 30% of your time charging, loading, infusing the data into the memory of the GPUs, it's not a very good use of your dollars. So I think it all has to do around latency.

Why does Theory Ventures not invest in AI hardware?

Matt Turck [11:10] All right, so let's get into AI proper. Maybe as a mental model to go through, let's think of it as a cake with diverse layers of the cake. Starting from the bottom of the cake, for each area we can talk about what you're excited about, less excited about, and where you think the opportunities might be. So at the very bottom, we're talking about GPUs, GPU infrastructure. Presumably, given the Theory Ventures model, which we'll talk about later, that's not an investable area.

Tomasz Tunguz [11:20] The hardware we'll stay out of. I think that's—

Matt Turck [11:24] Or the CoreWeaves of the world, or those businesses are growing really fast. Yes.

Tomasz Tunguz [11:45] But I think given our fund size and the capital intensity, that's beyond the scope. I think you really need to be pretty sophisticated when it comes to financial engineering, what kind of debt products you use. It's kind of like a REIT, real estate investment trust, in that you'll produce really good profits, but you need a lot of debt and you need special relationships with NVIDIA and the other GPU vendors to compete.

Matt Turck [11:47] That's the whole thing.

Tomasz Tunguz [11:48] Do you spend time there?

Matt Turck [11:53] No, for the same reasons. I do find it fascinating, though.

Tomasz Tunguz [11:54] Yeah.

Matt Turck [12:16] Especially the recent CoreWeave announcement. Talking about pivoting from crypto to AI, that was masterfully done, but it's great. It's actually sort of funny that this would be a business out of New Jersey. For sure. I don't know, maybe that's the next thing after the rise of the New York scene. You're going to have the rise of Jersey City.

Tomasz Tunguz [12:20] The PATH train startups.

Why do big companies invest in cloud infrastructure?

Matt Turck [12:38] Yeah, exactly. Awesome. All right, so that's a layer of the cake. And I think you were at the bottom of the cake, and you were, I think, writing in your great blog, giving some numbers around the CapEx investment in the cloud from the big vendors.

Tomasz Tunguz [12:59] $60 to $70 billion per quarter now from the top three clouds, and it just keeps going up. And you see Amazon spiked almost 18 months before Google and Microsoft did. So they must have seen some of these inference workloads coming. And then Google and Microsoft, yeah, $14, $15 billion a quarter. I think Microsoft just announced $3 billion in Wisconsin, and I have to imagine a lot of it will be in the data centers, but a lot of it will be in power.

Matt Turck [13:12] So is that good news or bad news for all of us that they are making those crazy huge investments?

Tomasz Tunguz [13:37] It's great news because Microsoft, in the last quarterly announcement, said that they're capacity constrained, right? So there are more AI workloads than they can support. And so for you and I, or for these startups that try to grow really fast, they need access to those clouds. And ideally, somebody else is paying. It's not venture capital dollars that are being used to buy and manage those GPUs. So it's a huge benefit. Right, yeah.

Matt Turck [13:54] All right, so going from the infra layer of the cake, or hardware layer of the cake, to the software layer, maybe let's start with foundation models. So, sort of the same questions. Like, those are huge kind of dollar plays. Is that something you're interested in?

Tomasz Tunguz [14:19] So we study it a lot because it matters. But, okay, so here's the thing, right? You look at Llama 3 8 billion, trained on 15 trillion tokens. I run it on my MacBook. It's awesome. Like, latency is super low. It works really well. It's basically on par with the Mistral 8x7B mixture-of-experts model. And I've stopped going to OpenAI because I can run it on my machine. It's faster, and it's integrated in my email client, and there are a bunch of advantages to it.

Tomasz Tunguz [14:51] And so I think the small language models, particularly for B2B applications, if they continue to improve their performance like this, make a lot more sense. We ran this analysis: GPT-3.5 Turbo and the cost per inference compared to the next most recent model, there's a 160x difference in pricing. And so what does that tell you? Well, a model that's six months out of date loses its pricing power rapidly, right? It commoditizes really fast. And then you have this open-source dynamic with Llama, where Meta and OpenAI are competing.

Tomasz Tunguz [15:25] So I think we'll see a lot of innovation there. I think we'll probably see a bifurcation of the model sizes, where the Llama 3 8 billion-parameter model, on an MMLU basis, which is the high school equivalency, is doing pretty well. I mean, it's pretty good. And then you have the 400 billion-parameter model that's still being trained. I don't know how long it will take. The question is, when do you use that super expensive, absolutely massive model? Is it a general search use case, or when do you need that amount of reasoning capacity and knowledge?

An investor view on the foundation models.

Matt Turck [15:47] Yeah. And I think you wrote something about this, if I remember correctly. Was it like the rise of kind of like hybrid architectures, where you use the big models for one thing, the small models for one thing?

Tomasz Tunguz [16:09] Yeah, so we call this constellation models. So the idea is a query comes in from a user, then there's a classifier, and it could be a generative classifier, it could be a classical model. It identifies the query and it says, okay, I've seen this query before. Let's say it's like a finance, a stock lookup or whatever. A query about an annual report of a company is probably a better example. Pass it to the small language model that's fine-tuned for that particular use case, like a Bloomberg LLM, let's say.

Tomasz Tunguz [16:43] Contrast that with a random query that the system has never seen before. It would route that to one of these bigger, general-purpose models because they can handle a much broader input distribution of inputs and produce a much broader distribution of outputs. On the other hand, those models are far more unpredictable and chaotic, and so you do want to kind of control them. But we think that the next-generation architectures look like that, at least for handling very simple tasks. What we're starting to dig into now is chaining is becoming the big thing, right?

Tomasz Tunguz [17:09] So how do I make an LLM actually act like an intern? The MMLU score says they're just like interns, or like high school interns. Now what we want to do is say, go and research whatever it is, fusion, and figure out what are the top five things that are happening in fusion. Well, if you hire an intern from high school and you tell them, go research fusion, I'll see you on Friday, the report that you get, I guarantee you, will be terrible because they'll be slightly off at the beginning and then they'll continue reinforcing, and all of a sudden they'll end up on Mars.

Matt Turck [17:20] Yes.

Tomasz Tunguz [17:41] And I think one of the challenges that will happen with some of these large language models, particularly when they're chained or we let them operate for hours at a time, is the intern effect, where they're off at the beginning. And that's just because they're chaotic and non-deterministic. So the output is just a little bit off of the first step, and then on the second step, the output's just a little bit off, and then the third step, the output's a little bit off.

Tomasz Tunguz [17:57] And so you multiply that error through a pipeline of 15 different steps. So the question then is, how do you mitigate that error? And so a couple of companies that we've met, they have a generative step, and then there's a classical model. So let's say we were researching fusion. The generative step could say, well, we could research this kind of fusion, that kind of fusion, this third kind of fusion.

Tomasz Tunguz [18:16] And then the classifier would say, research this one, and it would narrow down the range of acceptable outcomes. And at each step, it would quash the error that otherwise would compound.

Matt Turck [18:17] Yeah.

Is Gen AI going to replace traditional AI?

Tomasz Tunguz [18:36] This is, like, a very, very early mental model for how this could go. But I do think we need another technique, which would be actually using two adversarial—like one generative model and another that's adversarial next to each other—kind of question here. But I do think we'll need that, some sort of guardrails. I don't yet know if that's the implementation.

Matt Turck [19:12] So that ends up being query. Then you mentioned classifier. Some people talk about routers, that concept, then multiple models. And interestingly, you said generative AI, but also classical AI, which is sort of interesting, right? There's been that sort of trend in conversations and Twitter and whatever that generative AI is going to replace everything. But in reality, there's absolutely a huge room for classical AI that works on tabular data and all the things. Is that what you think? I agree.

Tomasz Tunguz [19:40] I think there's a role for both. The way I think about generative is, it's a really great knowledge compression engine, right? Like, I can compress the internet into a model that's, like, three gigs, right? Let's just say for Llama 8B. I can also use it as a classifier. Lots of startups use it. Email classification, for example, in order to identify phishing, which it's very good at. And then I can use it to create. The hard part is it can be really noisy and unpredictable.

Tomasz Tunguz [19:49] And so you need—you'll need both: time-series prediction and large language models.

Matt Turck [19:50] Yeah, maybe not.

Tomasz Tunguz [20:06] Yeah. And so it's an and. It's an and. And I think we're at a stage where this is a hammer and everything's a nail. And now we're going to start to understand, okay, how do we put together these constellations of models to get the ultimate underlying performance? I think the other really important part of all this, and something that we talked about earlier this week, was there's the model itself, and then there's the information retrieval.

Tomasz Tunguz [20:38] RAG is the rage now. And as the cost to train all these large language models becomes bigger and bigger, Amazon's talking about $1 billion for a single run or $10 billion for a single run. It's beyond the realm of startups to really spend any time there. So maybe the core intellectual property advance actually happens in the information retrieval step and the embeddings and the vectors, as opposed to the underlying models.

Does the Theory Ventures invest in AI tooling companies?

Matt Turck [21:05] Either you or Thierry, or both, maybe that's all the same thing, wrote a good piece on, what was it called? LLM search. Exactly right, like a few days ago. Maybe we'll add that to the show notes, as cool podcasters say. But that was a good piece. Okay. And from an investment perspective, so that layer of the cake, like the developer tooling, we started with models and we got into the tooling. Is that something where you have made investments recently?

Tomasz Tunguz [21:32] We have made two investments, two ex-Google teams. One is building analytics for large language models. So understanding: how are users engaging with an LLM? What are they talking about? How long are the session lengths? What's the quality of the conversation that can be used to inform evals, evaluations of the underlying model? And so it's really important for a marketer or a product manager to understand: is this robot actually doing what it's supposed to? The other one is another ex-Google team called Superlinked.

Tomasz Tunguz [21:45] They're trying to create a new category, which is called the vector computer. And the idea here is, if you've ever watched YouTube Shorts or TikTok and you watch three different videos, one on chess—

Matt Turck [21:46] That's all I watch.

Tomasz Tunguz [22:13] Yeah. So let's say you watch three different videos and the last one is on chess, and you spend a lot of time on that one, you will see more videos about chess. So there's a machine learning system that's updating in real time that's combining different kinds of data. It's combining qualitative information, the text of the videos, with quantitative information: how long you spent on the video, did you engage with the video? And each company will need to kind of—

Matt Turck [22:15] And that's providing that as a service?

Tomasz Tunguz [22:37] As a service. So the idea is each company will need to embed a different formula for defining vectors for the content or the recommendations that you need to make that will have many different components that will be both qualitative and quantitative. And so it's a service that allows you, as a product manager and engineer, to define: what do I want in that vector? Is it dwell time? Is it the text of the video? Is it the total number of likes?

Is investing in Cloud companies better than investing in AI-powered applications?

Tomasz Tunguz [22:54] Is it how much it spiked? Is it tweets? Is it whatever? So you can stuff a whole bunch of different information, and then that is updated in real time in your vector database. And so then you have much more performant machine learning systems. That's the idea of a vector computer.

Matt Turck [22:56] Applications. Is that interesting?

Tomasz Tunguz [23:00] AI-powered, AI-native. Yeah, it's every other year.

Matt Turck [23:03] Agentic AI applications.

Tomasz Tunguz [23:27] Yeah. Five trillion. In order to create the same level of market cap in the application layer, you would need about 100, the top 100 applications, like Netflix and Snowflake and all those guys. So there are many, many more opportunities at the application layer. I think the big question is, where do the incumbents play and where do the startups play? And I think one of the biggest areas initially was RPA. Our analysis shows that a lot of the existing robots, about 40% of the existing robots deployed in RPA, are broken.

Tomasz Tunguz [24:03] And that's either because the human process changed or the software underneath it changed. And GenAI helps you solve that brittleness with a bit of robustness. Now we're seeing sort of the unbundling of that, where it's much easier to create an automated SDR. It's much easier to create an automated UI designer or UI design. And so it's like use case by use case. We just invested in a company that's automating security analysis for the average enterprise of 75 security products. Each of those security products produces alerts, about 5,000 to 7,000 alerts per company per day.

Tomasz Tunguz [24:31] Less than 1% are reviewed. There's a lot of rote work. Customer support would be another one. And so the question is, okay, where can a startup play where there's no incumbent with a massive distribution and data advantage? And this is actually a much harder question to answer. The other dynamic within the application layer that we're trying to answer is quality of revenue. So when you have a fast-growing, fast-moving space like this, there can be a lot of demand in a very short period of time for a product.

Tomasz Tunguz [24:45] And the preferences of the buyer and the underlying technology will change also very fast.

Matt Turck [24:55] Yeah.

Tomasz Tunguz [25:13] So five years ago, if you and I were looking—I mean, and you've done this many times—you invest in a company, one to five to 15, let's say you had a very high degree of confidence that that business would be worth $500 million, maybe a billion, could be a public company, because the quality of that revenue is impeccable. It's gold standard. But in the AI ecosystem, we've seen many examples of companies where you have these massive rises in revenue and then all of a sudden either a plateau or an atrophy of that underlying revenue because the architecture has changed or whatever.

Matt Turck [25:38] Is that because the architecture changed, or because that was an innovation experimental budget in the first place anyway, and people needed to do AI immediately and just jumped on something?

Tomasz Tunguz [25:55] Yeah, this is a really good point. So I think there are a lot of buyers who are just trying to understand. Nobody wants to be outmoded. Everyone wants to have an answer when their manager or their CEO or the board says, "Hey, what do you think of this AI thing?" You want to say, "Oh, I'm at this great company and we're just starting this pilot." And then, I mean, we know, right, machine learning systems: it's easy to get to an 80% solution, but all the value is in that marginal 15 or 18 to get to 95 to 98% accuracy.

Matt Turck [26:05] It's very hard.

Tomasz Tunguz [26:12] Yeah. And so I think you're right. There are a lot of trial budgets that are happening. But then the underlying architectures are changing too.

Matt Turck [26:12] Yeah.

Copilot AI vs full-execution AI.

Tomasz Tunguz [26:40] RAG is relatively new, and I think they will change a lot. I was on a panel recently where we were talking about large language models and the parallels to search engines, where you may not necessarily want to be the first search engine because the underlying information retrieval architecture could change pretty phenomenally from an Oracle database to vanilla box servers. And that cost advantage from AltaVista to Google makes a pretty significant difference in the ultimate attractiveness of the business.

Matt Turck [26:54] In terms of applications, the now good old question: copilot mode versus full execution mode. Do you have any preference? How do you see things evolving?

Tomasz Tunguz [27:22] We're not religious about it. I think coding, I think you'll have both, right? You'll be a software engineer, use Copilot or any one of the others. And let's say you are reviewing a pull request, very likely that you're more likely to use a copilot than you will use a full agent. Whereas if you look at Devin and you want to create a stub of a Shopify store, very likely that you'll end up using full automation there. So we think there's room for both.

Tomasz Tunguz [27:47] There's a difference in the productivity gains. So Microsoft and ServiceNow have said there's a 50% to 75% increase in developer productivity. I don't know if they're measuring that through lines of code or some other composite metric. There's no data yet on how productive agents are, and it's probably use-case specific. The crudest analogy that we have, which is probably the best that we've found, is for a mechanical robot on an assembly line: how many humans' jobs does that robot replace?

Tomasz Tunguz [28:25] Five. So theoretically, if it's anywhere close, an agent is probably four to five times more productivity-boosting than a copilot. But I think it'll be really use-case specific. I couldn't imagine an agent in the legal domain or writing—maybe writing an NDA would be okay—but putting together a bespoke term sheet or an auditor being replaced by AI. Yeah, term sheets.

Matt Turck [28:35] That's a highly complex legal document where all our VC expertise over decades gets embodied.

Tomasz Tunguz [28:36] Come on, co-sale rights are tricky.

A case for specialized LLMs.

Matt Turck [28:59] And a little bit to that vertical kind of use case point and the earlier part about small specialized language models. Is that an area that you find interesting, where you have, I don't know, LLMs for bio something, or LLMs for supply chain?

Tomasz Tunguz [29:28] Yeah, they will exist. I think you'll probably see generic clouds that offer model families that are not customized. We're starting to see some companies that are either vertically integrating the stack from the GPU and networking layer all the way through the application, particularly images and video. The cost advantages are significant. There are yet other clouds that are specializing: here's a collection of small language models that are focused on finance or legal use cases. Again, to kind of bound—

Tomasz Tunguz [29:53] I think the question there is, aside from the financial engineering, what's the sustainable competitive advantage? There's definitely a case to be made that there'll be a DigitalOcean-like business, where catering to different needs of developers rather than a broad, super-complex cloud makes sense. I don't know. We have a hard time answering that sustainable competitive differentiation question at that layer.

Gross margins in Gen AI: is it profitable?

Matt Turck [30:09] And you wrote some very interesting content on the cost aspect of this and how people should be thinking about cost, both from a customer perspective, but also from a startup perspective, like your sort of gross margin kind of thing.

Tomasz Tunguz [30:36] Yeah. So this is the question we were asking. The average publicly traded software company has 71%, 72% gross margin. Pretty good business. So then the question is, okay, throw Gen AI in there. What happens to gross margins? Well, Google, when they initially released Search, their generative search experience, the cost to serve a query was 10x a standard one. Now, they'd had 20 years of working on reducing the cost of a standard query. So it's not apples to apples, but it's still significantly more expensive.

Tomasz Tunguz [31:02] And so you definitely have an increase in costs on the hardware side. The question is then, with an automated SDR, let's say, and then automated marketing or Copilot-assisted engineering, are the efficiency gains there enough to offset the marginal cost in your inference? And think, like, we ran a little thought experiment, and the answer is yes. I mean, and these are just guesses, but we think the average company in 18 to 24 months should have 2 to 5 percentage points more gross profit.

Tomasz Tunguz [31:34] And if you look at the S&P, we actually think—I'm not a public markets guy—but I actually think the earnings that will come from many major businesses will have a one-time massive jump as a result of AI. You look at Klarna cutting two-thirds of customer support, and we can debate the quality of the answers, but they did it.

Matt Turck [31:34] Yes.

Tomasz Tunguz [31:59] And all of those dollars are flowing to the net income line. I was chatting with the CEO or the founder of a publicly traded CRM company, and his estimate was that by the end of 2025 to 2026, two-thirds of SDR positions will no longer exist. And so I don't know if that's true. And the reality is those people, they will be repurposed for other things, probably managing the robots as opposed to actually doing the work—a more enjoyable job. But if that's anywhere close to true, there should be a lot of efficiencies that should accrue to the bottom line for a lot of these software companies.

Tomasz Tunguz [32:19] And if that's the case, you and I both know, then the magic number of the sales efficiency of this business goes up a lot, which means that their willingness to spend on customer acquisition should also go up.

Matt Turck [32:20] Right.

Modern Data Stack: is it still a thing to invest in?

Tomasz Tunguz [32:35] And then over time, does that advantage sort of atrophy away because everybody's just spending more on Google Ads or LinkedIn or events or whatever it is? I think you'll see a pretty significant increase in short-term profitability for a lot of these companies.

Matt Turck [33:09] Modern data stack, the hottest area of 2021. I was actually on this podcast, had a good chat with Tristan at dbt around the general topic of, is the modern data stack dead? Which was interesting, especially given all the impact that Tristan has had in terms of coining the term and helping build the ecosystem in the first place. So where do you stand on this? You, over the years, made some great investments in the space. I think you continue to do that.

Matt Turck [33:15] Do you continue investing in it right now?

Tomasz Tunguz [33:44] We are. It's very different than it was two years ago. So two years ago, whatever, Snowflake, 171% net dollar retention, everyone was spending. Then the Fed raises rates, and cost-cutting becomes the number one priority for all these data leaders. I think that's now run its course. The other dynamic that's really important in the modern data stack is, as your landscape shows, it's grown enormously. And so there's a lot of fragmentation of budget. And the dominant sentiment that we hear from many data buyers is, please don't sell me another tool.

Tomasz Tunguz [34:22] I don't need another tool. So that's definitely happening. Then you have the Iceberg separation of compute and storage, where the biggest companies really want to keep their data on their own systems and then use Iceberg and Parquet and Arrow as a specific format, and then have query engines on top. And we've been talking to several major banks where they have four query engines running on top of these Iceberg data lakes. That's somewhat of a deflationary force within the ecosystem. About 11% of Snowflake's revenue is markup on storage.

Tomasz Tunguz [34:54] So it's not huge, but it is important. And so the question is, okay, let's reimagine the modern data ecosystem, first at the data engine layer, the database layer, where there are all these different query engines. And each time you have a workload, you're actually picking the right query engine for that workload. That's what's happening now in the most innovative companies. And so that changes the architecture from the database up. There are questions about, okay, as you move up the stack, as you move into the data transformation layer within the database, let's call it like dbt, or you move into the ingestion layer, what happens there?

Tomasz Tunguz [35:37] What also happens? What changes in the BI? I think at a high level right now, there's just a lot of rationalization of spend. I think we'll see a pretty significant wave of consolidation and simplification. Snowflake and Databricks is the dominant battle here. And Snowflake, clearly with Sridhar and with Baris running AI, you can see the pace of innovation coming out of that business now is real, and they're pushing to broaden their ambitions a bit. So I think you'll see quite a lot of consolidation in the next couple of quarters.

Matt Turck [35:47] Companies going away or companies getting acquired?

Tomasz Tunguz [35:48] Companies being acquired.

Matt Turck [35:50] Any private-on-private merger kind of thing? Yeah.

Tomasz Tunguz [36:02] So I think you'll see Snowflake and Databricks being really acquisitive. That's not based on anything. I mean, it's just my hunch. But Google, Microsoft—

Matt Turck [36:07] No public information was discussed in the context of this podcast.

Tomasz Tunguz [36:36] But you look at Amazon, Google, Facebook, Microsoft, all under antitrust, basically impossible for them to acquire. You don't really have the big checkbooks at the table anymore, which means if you're an acquirer like Snowflake or Databricks, you should have some degree of pricing power. And then both of those companies are going to be looking to compete to offer end-to-end data platforms for their customers. And so Snowflake is stronger in structured data, Databricks is stronger in unstructured data. Each of those has roadmaps that are combining them.

Tomasz Tunguz [37:01] And so how do you accelerate that? Well, you acquire really great teams. And so that'll be a really big dynamic that will drive. And then if there's a change in policy at the FTC along the antitrust lines, then I think you'll see even more because the clouds themselves, you look at Redshift, in the buyer's mind, has really fallen behind in the next-generation cloud data warehouse.

Microsoft Fabric and its impact on the market.

Matt Turck [37:08] What do you make of Microsoft Fabric? Is that crashing the party?

Tomasz Tunguz [37:14] So there was a stat, I think it was 11,000 customers.

Matt Turck [37:14] Yeah.

Tomasz Tunguz [37:37] And I can't tell if that's existing customers that were on similar products, Power BI, just moved them in this category. That said, let's look at that number. We'll discount it. It is a really important force in the ecosystem. Microsoft's ability to cross-sell is second to none. You look at Teams and Zoom, that whole dynamic. I think something very similar will happen to data. I think overall, there's a wistfulness for what could have been with Databricks within Microsoft that they would like to replicate and bring back in.

Tomasz Tunguz [38:13] And so they've restructured. My understanding is they've restructured the database team very recently in order to enable Fabric. They've been executing phenomenally well there. Our market analysis suggests that customers are very happy with Fabric and it's working. And so again, going back to that, if you're Snowflake or Databricks or a smaller player in the ecosystem and you need to compete with that broad panoply of solutions that exist within Fabric, you'll need a bunch of different point solutions.

Matt Turck [38:34] In particular, an AI layer, right? For both of them. I guess Databricks is originally coming from the machine learning and AI world, went into the big kind of storage database game, and is now going back, adding all the things because it feels like the ultimate game is data plus AI.

Tomasz's thought on Motherduck and DuckDB.

Tomasz Tunguz [38:50] Totally. Yeah. Well, Databricks SQL, the product, is $200 million in revenue, doubling. And so the serverless—it's actually the serverless part of the business that's growing the fastest. And two years ago, I couldn't have told you I knew of a single customer using that infrastructure for BI workloads. But now that's very different.

Matt Turck [38:56] Where do Motherduck and DuckDB fall in that whole dynamic?

Tomasz Tunguz [39:22] Yeah, that's a really exciting business. I think it's a query engine. So it's one of those query engines. We see it more and more. There's a couple of different dynamics that are really important for that business. The first is there's broad open-source traction. It's really easy to try. And so instead of spinning up a Snowflake instance, you start on your computer. The second dynamic that's really important is this idea of hybrid execution. So I start locally.

Tomasz Tunguz [39:29] Jordan, who's the founder, was a GCP tech lead. Eighty percent of those workloads are less than—

Matt Turck [39:29] BigQuery?

Tomasz Tunguz [39:51] BigQuery, yeah, sorry. Eighty percent of those workloads are less than 100 megs in size. And so you can do that on your machine. You can do like 10 gigs even with a really fancy machine. So can you start on your computer and then burst to the cloud as necessary? That's one really important path for them. The other is these hybrid applications. So Omni and Hex both use Motherduck, and Motherduck, the cloud, will send 1 to 2 gigabytes of data down to the browser.

Tomasz Tunguz [40:15] It will be stored in memory, and then all of the additional data processing is done inside the browser using WASM, and it's compiled to be really fast. And so you can manipulate literally a gig of data and have the UI update in 100 milliseconds.

Matt Turck [40:15] Yeah.

Tomasz Tunguz [40:33] So it's better even than Tableau in terms of data scale. So that's a really powerful new model. And then we just announced pricing, and serverless is significantly less expensive than other forms of serving applications. So I'm really excited about where the flock goes.

Where do BI tools fit in the Modern Data Stack?

Matt Turck [41:08] Very good. And you're also an investor in Omni, actually, interestingly, both from your time at Redpoint and now with your Theory Ventures hat on, which is really interesting to me because BI, of all parts of the modern data stack, is really, to me, the one that sort of felt like the kind of unloved child where you've seen less innovation. So what's the story about Omni, and where do you think, I guess, modern BI should be going?

Tomasz Tunguz [41:33] Yeah, so BI, $16 billion category. You're right, it's a super competitive, hard-to-differentiate category. I think it's broadly misunderstood. But the way that we think about it is it has swung between a pendulum of control to empowerment. So during the 2000s era, there were four centralized BI companies: MicroStrategy, Cognos, Business Objects, and Hyperion. And then Tableau came and unbundled the visualization layer. So it swung from really tightly controlled, centralized control to anybody can do whatever they want with their data.

Tomasz Tunguz [42:03] Then the cloud data warehouses came, and Looker said, well, let's go back to centralization with the data modeling. And what Omni is trying to do is narrow those swings and say you can have the control and the modern data model, and an individual marketing analyst who wants to define a particular kind of cost of customer acquisition can use an Excel-like UI to do that. And if the metric is awesome, they can say, I want to promote this to the team, to the group, or to the entire company, and that then folds into the underlying data model.

Tomasz Tunguz [42:23] So the data team can define metrics, but also an individual person in the company can define metrics, and there are approval flows to marry the two.

Matt Turck [42:31] So they're unbundling, is that right? Yeah, I lost track of the—in general, in software, the bundling, unbundling, rebundling, re-unbundling.

Tomasz Tunguz [42:59] Yeah, it's sort of a bundling. I mean, you can think about it as—so one of the team demonstrated it to one person in the ecosystem, and they said it basically has the data model of Looker and the flexibility of Tableau and the spreadsheet interface of Sigma. So you can think about it as a bundling. That's exactly right. And then underpinning it, I think one of the things that's misunderstood about BI is it's really a governance tool. The charts themselves are important and interesting, but the surface area of the products is enormous and takes years to build because it's all about how do we make sure that the data is going to the right person, it's correct.

Tomasz Tunguz [43:41] That's the value proposition of BI. A lot of people focus on the charts and whatever, the choropleth maps and that kind of stuff, but it's all about permissions. The other part about BI that is misunderstood is that embedded BI, BI within products, is an absolutely enormous market. At Looker, when we sold the business, a third of the revenue was embedded. And out of a team of about 700 to 750, there were two people working on that product. So there's really—everybody needs analytics.

Tomasz Tunguz [44:01] It's all changing. If you can have internal BI where a CEO is looking at the revenue by product, by geography, and the same system is powering your trucking system, there's a lot of advantage because there's a single data model and you don't have these islands of data fragmentation.

Why has the democratization of BI never happened?

Matt Turck [44:34] To double-click on your governance point, indeed, the reality still today of BI is that you have a handful of analysts who have access to the BI tool, and then the rest of the organization asks those two analysts for information. And then there's basically a pecking order where, if you're the CEO, you'll get the data very quickly. If you're a middle manager, take a number. You may get your reply in a week or two weeks. Why has the democratization of BI, which is an old dream, never happened?

Tomasz Tunguz [45:06] So there are two schools of thought within data leaders. The one school of thought, what would you call classical, is that it's better for the organization to have less data if the data is all correct. And it's very easy to misinterpret data. There needs to be a lot of education around statistical significance or how to design experiments or how to process data. And so I would say that's the majority of the industry. There's a minority of the industry who says the only way forward is for data teams to become like product teams, where they build platforms of tooling that allow individuals to ask and answer their own questions.

Tomasz Tunguz [45:44] And there's a huge education component. So they have office hours. Warby Parker started this, I don't know, 10 years ago. ZoomInfo is really good at this. There are other companies, but there is a divide within the ecosystem about which is the more effective path. And there's no—it's philosophical right now. There's no—it's kind of like, do you put SDRs in the marketing team or in the sales team? All roads sort of lead to Rome, and you could probably end up at the same place, but stylistically there are different preferences.

How do acquisitions happen? Can you engineer them?

Matt Turck [46:24] So we talked about consolidation. We talked about Looker a little bit, $7 billion acquisition by Google. Anything you've learned for all the founders who may be listening right now that may be part of the group of companies that you said should be acquired by Databricks or Snowflake? How do acquisitions happen? Is that something you can engineer? Is that something that you need to position yourself for? What have you seen and learned?

Tomasz Tunguz [46:57] Yeah, so the answer is someone's putting their career on the line. So when someone buys a business, not an acqui-hire, but let's say, like, $50 million to $50 million-plus, someone in the acquiring company is saying, "I am betting my career, effectively, that we should buy this company at this price." And that doesn't happen overnight, right? It doesn't happen in the course of, like, two or three months. It probably happens over the course of 12 to 18 months.

Tomasz Tunguz [47:22] It's a big, long enterprise sale. You can think about it that way. And there, it is a multi-party sale where typically you have a GM or a head of product who decides, "I really need this business as part of the product portfolio." And rather than building it, which would be much easier politically to do, to navigate that budget, I think we should buy. And I'm willing to make a case and go first to my manager, then to my VP, then to the board to justify it.

Tomasz Tunguz [47:58] Involve legal, corp dev, and many other teams, and manage what must be a very difficult cross-functional effort to get this through. And so somebody really has to care. There has to be a very strong personal relationship between the founder of a business and that buyer, because that buyer is putting together a three- to five-year plan to demonstrate some positive ROI on that acquisition. So they'll take the startup's plan and then discount it pretty significantly, and need to justify it somehow.

Tomasz Tunguz [48:21] So shotgun acquisitions very rarely happen. So that's one of the things that I've learned. It's all about that product relationship or the GM relationship. So initially it starts with—and this is why, if a potential acquirer asks to meet, you should really always take the meeting. Don't spend a lot of time, but start building the relationship. And the thing that surprised me at the beginning in venture was even if you're competing, you should be friends because you have a lot more to gain from sharing those experiences and understanding of the world than you do from isolating yourself in the ecosystem.

Tomasz Tunguz [48:53] And so it's the same with a potential acquirer. If you're the CEO of a company and there's a big incumbent, you should know the management team of the competing product. And so anyway, that's like one life lesson, or one thing that I've learned.

Key ingredients to build data infrastructure business.

Matt Turck [49:15] And in the same vein, building data businesses, data infrastructure businesses, anything in terms of the pace, the dos, the don'ts, the positioning, any kind of pattern you've seen around successful data businesses or data infra businesses?

Tomasz Tunguz [49:36] So I think data is a much more ecosystem-driven space than a lot of other spaces. And what I mean by that is, if you're a database, you need the upstream ETL vendors, you need the dbts and the Zubikos in the center, and then you need the BI vendors. When Looker went to market, we initially partnered with Snowflake, and we were bringing Snowflake into deals, and that worked really well. And then Snowflake started to grow really fast, and Snowflake was bringing Looker into deals.

Tomasz Tunguz [50:05] And so why is this? Well, data is an ecosystem where there are lots of individual point products, but what a buyer wants is end-to-end. They want me to solve the problem. So that Looker-Snowflake combination was, I need a reporting stack, I need this report, and it needs to be this fast and be shipped on this day, and it needs to be in this format, delivered to these people. Well, you need a database plus BI. And so figuring out what is that end-to-end stack that you're selling and becoming part of that stack is really critical.

Tomasz Tunguz [50:32] It really helps quite a lot. And the way that it starts is it's typically AE to AE. So a Looker AE plus a Snowflake AE get together and they start selling two or three contracts. And then friends of that account executive say, "Hey, Matt, how'd you hit your number last quarter?" And they say, "Well, I happen to be selling this Looker thing, and it's working." And so the usage is higher, and then that goes up to the team, and then that goes up to the director, and then it gets a lot more serious.

Tomasz is a founder now! How does it feel?

Matt Turck [50:41] So you're a founder now?

Tomasz Tunguz [50:42] Yeah.

Matt Turck [51:01] How does that feel? Maybe as a last theme in this conversation, talk about your evolution from Redpoint to Theory. So I guess, why did you decide to start your firm? And how does it feel? How long has it been, a couple of years?

Tomasz Tunguz [51:04] It has been since September '22.

Matt Turck [51:05] September '22.

Tomasz Tunguz [51:06] Yeah, 18 months.

Matt Turck [51:08] 18 months. So how do the first 18 months feel?

Tomasz Tunguz [51:29] It's exhilarating. I have a lot of—I mean, I started a company when I was 17. It was a little bit of a toy business, but starting it has been a real gift and a privilege. I think the first part about it that I really enjoy is setting out a strategy for the very long term. So we have a 20-year plan inside of the firm. The second is there's this moment when—and I imagine it's like when a founder receives their first check—somebody believes in you.

Tomasz Tunguz [52:05] The first commitment into the fund, or the first hire, or the first 10 people that join who believe—that's really special. And I don't know if I appreciated that to the extent that I should have, being involved in the startup ecosystem. I think the other thing that I really love about the firm is we're able to move really fast. And we just had an offsite on Monday. Everybody inside the firm is responsible for one experiment per year. We expect 70% of experiments to fail.

Tomasz Tunguz [52:17] So we have programs like this where we're just trying to innovate and push, and being able to architect those programs and execute them is exhilarating.

Matt Turck [52:20] What's an experiment? Is there one per function kind of thing?

Tomasz Tunguz [52:43] Yeah, it's typically one per person. So we have this new format with—we have about 100 theorists, as we call them, who are experts around the firm. And we have this new format called quartets, where we bring two or three of them together and they talk about stuff. And the goal is just to build community. And that's been really well received. So that's a really small example of that. I think other things that we're doing are more sophisticated modeling around portfolio construction, which is a really big emphasis for us.

Tomasz Tunguz [52:53] And thinking about active resource management.

Matt Turck [52:57] Interesting. And the fund is $230 million?

Tomasz Tunguz [52:59] Yeah, $238 million.

Matt Turck [53:02] Exactly. $238 million, investing 1 to 20.

Tomasz Tunguz [53:03] Exactly, yeah.

Matt Turck [53:07] So that's seed to B?

Tomasz Tunguz [53:14] Yes, mostly seed and A. So, yeah, exactly. Omni was a B.

Talking numbers: Theory Ventures' financial model.

Matt Turck [53:22] Okay. And then you mentioned being concentrated. What does that mean? How many portfolio companies per fund?

Tomasz Tunguz [53:23] 12 to 15.

Matt Turck [53:38] 12 to 15. Okay, so that is a fair amount of concentration. Okay. And then what is the value proposition to founders? Is that your deep knowledge of the space, or how do you position?

Tomasz Tunguz [53:40] That's right.

Matt Turck [53:41] Yeah.

Tomasz Tunguz [54:08] So we lead with expertise. The goal is really to understand the space. So our CEO is amazing. Her name is Lauren. She was building data architectures at Palantir for healthcare, for many of the largest healthcare companies in the world. And so she joined, and we have an intelligence team that is four people today that is focused on active research. So we'll research spaces for months or quarters before, and we try to preempt. So about three-quarters of the companies we've invested in so far are preemptions.

Tomasz Tunguz [54:15] That's the motion that we're trying.

Matt Turck [54:31] Well, it's been a wonderful conversation. Sort of feels like—I know that sounds like a trite cliché, but very true in this case—it sort of feels like we could go another two or three hours. So yeah, I really appreciate it. Super interesting. And thanks for doing this.

Tomasz Tunguz [54:32] My pleasure. Thank you.

Matt Turck [54:53] Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.