MAD Podcast
    MAD Podcast

    The MAD Podcast with Matt Turck

    What Happens When Billions of AI Agents Hit Your Database? (Andy Pavlo)

    Andy Pavlo is the VP of Database Research at ClickHouse. We cover why agents could drive a 10–100x increase in database queries, why agents need sandboxed branches and permissions to avoid deleting production data, and how LLMs can reach roughly 85% of database-tuning quality in 15 minutes.

    10/08/2026

    Hosted by Matt Turck · with Andy Pavlo, VP of Database Research at ClickHouse

    AI agentsDatabasesClickHouseLLM tuningText-to-SQL
    Listen now
    YouTubeApple PodcastsSpotify
    1h 17m · 24 chapters
    Contents

    Transcript

    From vector databases and RAG to the age of agents

    1:21
    Matt Turck0:52

    I'm Matt Turck, and this is The MAD Podcast. Andy, welcome.

    Andy Pavlo0:53

    Hey, Matt, how are you doing, man?

    Matt Turck1:01

    Good. I'm excited to talk about all things AI, agents, databases, Larry Ellison. We have to. And then Wu-Tang Clan.

    Andy Pavlo1:12

    We never thought Wu-Tang Clan. Do we have to talk about Larry Ellison? I mean, it's a little early to come out of the gate, although he is hilarious because he's like the supervillain of databases. He owns a Hawaiian island.

    Matt Turck1:18

    Well, that seems to be working out for him. Yes. All right, well, we can graduate to talking about Larry Ellison later in the conversation.

    Andy Pavlo1:22

    I will say my contract with ClickHouse does allow me to talk about Larry Ellison. I made sure that was in writing.

    Matt Turck1:48

    All right, awesome. I want to take it from the top. So, two years ago, the topic of databases and AI came up. There was a whole era of vector databases and this whole idea that you need to do RAG in connection with chatbot searches. That was kind of two years ago. I think a lot of people's mental models kind of stopped there in terms of this intersection of databases and AI. Clearly, we are in the age of agents now.

    Matt Turck1:54

    How is that different? What does that change?

    Andy Pavlo2:17

    Vector databases, although you're right that, in terms of being in the zeitgeist of AI and databases, probably two years ago is when RAG and vector databases came into play. But the vector databases themselves, a lot of them were around since 2016. It's just now that the use for them became obvious because you can use them to enhance and improve agents doing things. So I think the vector database market is interesting because they were early in the space, but they were really pushing the idea that you want to do this approximate similarity search on vectors or embeddings instead of doing an exact search that database systems have traditionally done.

    Andy Pavlo2:58

    And this helps, too, when you're dealing with natural language, where the semantics of maybe what you're trying to look for or query against, you can't precisely define. Like, find me all the albums from artists in New York City that talk about doing certain kinds of activities. Right? It's very hard to define exactly what that is. And the magic of approximate nearest neighbor search, or these vector embedding searches in a vector database, is that you can look up things based on the semantics of the thing you're looking for, and it just sort of works out that way because you're not trying to do exact keyword searches.

    Andy Pavlo3:40

    So when ChatGPT blew up in 2023, when everyone became aware that this is now a technology that exists, the vector databases were in the right position at the right time because they could augment and improve these chatbots. But the moat, if you will, wasn't that big because at the end of the day, it's just an index. And within a year, pretty much every single database vendor had their own vector index. And now it's become sort of standard table stakes.

    Neon's stat: agents create 80% of databases?

    4:09
    Andy Pavlo4:09

    Every database system that we're using has its own vector index. Now, it doesn't say that the vector databases can't do certain things better than what the traditional data system could do. It's just now that everyone has that basic functionality. So I think that you still use vector searches to enhance and improve the chatbots or the agents. Now, you don't need to have a specialized one database that can only do that.

    Matt Turck4:31

    And agents now create databases, use databases, can make changes to databases. So maybe a year ago, when Databricks acquired Neon, there was this stat that I believe said that the number or the percentage of agents creating Neon databases went from 30% to 80%.

    Andy Pavlo4:36

    Or the databases created by agents was larger than everything else.

    Matt Turck4:43

    Unpack that for us. What does that mean, an agent creating a database? Why does an agent need to create a database in the first place?

    Andy Pavlo5:03

    The creation of a database. And then there's branching of an existing database. And I think the Databricks metric was about the branches being created by agents. So what does branching mean? Branching is when you have a database and you want to make a sort of copy-on-write snapshot of that database and have it look like two unique databases. And the reason why you want to do that is you could have your production database that's running your business, your application, where all your system-of-record changes are going into.

    Andy Pavlo5:35

    But then you want to have either a human developer or an agent be able to start trying out new things, make changes to the application, and you want them to operate on the real database. Because oftentimes, traditionally, you would make copies of the database or have a smaller version of the database, like a sample of what already exists there, because it's just too expensive to copy things around. So you have, like, a staging or development database, and then you would hopefully apply your changes to the production database, and you hope that once you go to the larger scale, everything works out.

    Andy Pavlo6:07

    So what branching allows you to do is it allows you to make a snapshot of the real database, have all the changes that you make go to this snapshot, and not affect the production database. And once you're okay with the changes you've made, maybe not necessarily for the data, but usually for the schema or the application code, once you know you've vetted it on the snapshot, then you can put it live in the production database. So now, with these agents, people are spinning up these coding agents to write new features, write new parts of their application.

    Andy Pavlo6:32

    It's not something humans are doing anymore. And now you want to have these agents vet their changes on a copy of the database, so you can create branches on these things. Often it just shows up when you're running testing as well. A lot of the time, the branches are created by PRs running in GitHub Actions or testing things like that. So you want to run your test code on a copy of the database rather than the live one.

    Andy Pavlo6:51

    So this branching capability is mostly being used for development. People are using agents to write a lot of new code, and so you want them to run in their own sandbox. You make a snapshot of the database. So the branching capability is super useful for that.

    Matt Turck7:03

    Yeah. At an even simpler level, an agent needs to create a database because agents take actions on the world. And if you ask an agent to create an application, an application at the end of the day is still a database.

    Andy Pavlo7:21

    Yeah. So certainly, if you say, "Create a database from scratch," then yes, your coding agent—you give it the prompt, say, "A billing application does this"—then yes, it will create a database right away. But most of the branches or new databases, if you want to call it that, are being created for this development. But yes, the agents can now create databases. For better or worse, you can have them drop databases. People often get in trouble with this because it's doing things that they shouldn't be doing.

    Why agents keep deleting production databases

    7:25
    Matt Turck7:28

    Yeah, there was last year as well.

    Andy Pavlo7:30

    Almost every year, there's some story.

    Matt Turck7:59

    There was some incident that was very interesting and got pretty widely publicized on X, for what it's worth. So there's Jason Lemkin that was trying to build something with Replit, and I think the Replit agent went into the database and deleted a lot of files, which then subsequently Replit made the right adjustments. Obviously, all of this is the Wild West, and people are learning as they go, and mistakes will be made. But what—

    Guardrails: the toddler-and-stairs rule

    8:19
    Andy Pavlo8:19

    It's the Wild West for people that don't know the history of databases. For years, we've had the capabilities to make sure people don't do things they shouldn't be doing, stupid things like dropping a data table or deleting records. Those controls exist. It's just everyone wants to reinvent the wheel, and people learn the hard way of just doing what's already been done in the past.

    Matt Turck8:30

    Let's expand on that point. So what does that mean? What needs to exist in the agentic world for databases that is not currently in place?

    Andy Pavlo8:51

    Guardrails, right? It's like you wouldn't have a new child, like a new toddler, and just not put up the guardrails so they don't fall down the stairs or put their finger in the socket, right? It's the same kind of thing for these agents. Think of them as toddlers, a little more sophisticated than that, but you don't want them to run wild or do whatever they want on a database. Now you could sit and babysit it and just say, "Approve, approve, approve," like in Claude Code or whatever.

    Andy Pavlo9:14

    But of course then you can't scale and have hundreds of agents trying to do things at the same time. So you want to set up the right permissions and controls so that they can't go into the production database and drop tables that they shouldn't be. And likewise, you don't want them to go read things that maybe you necessarily wouldn't want them to read as well. So if you have the right guardrails in place, and through branching or whatever means you want to create these sandboxes, you can avoid disasters. Like every month there's a story where someone lost their production database because the agent deleted something.

    Andy Pavlo9:41

    Did something they shouldn't do. With the right guardrails in place, you can assure that these things don't happen. People get lazy, they use the same password across the entire company or organization, and then they give the same password to the agent, and now wonder why the agent started doing things they shouldn't do.

    Matt Turck9:46

    Should those be the same guardrails as what currently exists?

    Andy Pavlo10:12

    Yeah, so maybe the larger question is: do agents, at least now, behave differently than humans? Other than just being able to run 24/7, I haven't seen anything that would suggest the fact that they don't develop software the same way or interact with data the same way that humans do. Right? At the end of the day, the data system only sort of exposes so many different capabilities, and that's the API that gets exposed to you. So as long as you can constrain those things for humans using those APIs, you could do the same thing for agents. But the level of sophistication of what those guardrails could be beyond just, like, don't run delete or update on a table.

    The four eras of database volume: 10–100x more queries

    10:51
    Andy Pavlo10:52

    The level of sophistication really matters on what sort of product you're using. And I will say this is something where I like open-source databases, but the enterprise commercial legacy databases, if you will, they've already done all this for humans in the past. So they have a lot of the protection mechanisms you would want. Just expose that to agents as well.

    Matt Turck11:04

    Understood. Is there anything that's fundamentally different compared to humans? I think you alluded to volume and scalability. Does that change everything, the fact that you have possibly billions of agents taking actions?

    Andy Pavlo11:24

    And how they interact with databases? I got to be honest, I don't know the answer. It's still early to say. I mean, the volume, of course, because agents are running 24/7 in a way that humans cannot. But we've seen similar patterns in databases before where we've seen increases in volume. And I would sort of say there's been three eras, and we're now in the fourth one. The first one was at the very beginning of just having a database in general, that you can now start putting in information and run transactions and store things for your organization in a way you couldn't before when it was all pencil and paper.

    Andy Pavlo11:58

    So probably like the '70s and '80s, there was an obvious increase in the volume of traffic you needed to handle in a database. But in that era, it was mostly humans interacting with the database at their jobs, because no one would go home and no one had a computer necessarily in the '80s, certainly not the level we have now. So that was sort of, say, one era where we had an increase in volume. Then in the 1990s, early 2000s, the internet comes along, and now everyone can interact with databases through websites and through the internet.

    Andy Pavlo12:22

    At home. So you saw a sort of spike in volume at that level. Then the next era would be mobile phones. And now everyone at any given time has various apps running. They can interact with websites, then put things and read things from databases. So I would say that would be the third era. So now agents, I think, is the next chapter in the story, where now every human not only has a cell phone, but they also can have maybe hundreds of agents interacting on their behalf that are then interacting, reading, writing data from a database.

    Andy Pavlo12:51

    At the end of the day, when you name an application, it's always going to be backed by a database. So it's going to cause the application to read and write information, read and write data. So I don't think we're there yet, because every human obviously doesn't have hundreds of agents running 24/7. But what's interesting about what could happen is, when you think of a human interacting with their cell phone or their laptop on a website or an application, they're kind of really only doing it when they're awake, right?

    Andy Pavlo13:24

    You can only do one thing at a time. Like, I can't sit with my one cell phone in one hand and a laptop in the other hand and do things at the same time. You're kind of only limited to what the human can actually do. But you can imagine a world where these agents are just 24/7, always running and acting as if they were humans interacting with things. So the volume would increase quite significantly. And then furthermore, unlike a web app, they're usually going to be— you could argue, how's this agent really different than IoT devices or cell phones now?

    Andy Pavlo13:56

    Like IoT devices, you're just streaming inserts of, here's the measurements I'm getting. With cell phone apps, you're usually reading data, right? But you can imagine these agents do more complex operations where it is a combination of reads and writes, and then that makes the traffic more interesting. So I think there's going to be potentially a 10 to 100x increase in volume and the number of queries that are going to hit databases. And I don't think we're there at that level yet, but I definitely have heard from my friends in industry that they're seeing the sort of thing we've seen.

    Andy Pavlo14:14

    We've seen the same thing at ClickHouse. What I don't know yet, though, is whether the sophistication of the interactions are more sophisticated or less than what humans can currently do. I've seen some people say, well, the type of queries that the agent would generate, or SQL queries the agent would generate, is actually more shallow because it's just trying to maybe read one or two tables at a time, try to figure out what's in there, versus if someone uses a dashboard or a BI tool like Tableau or MicroStrategy, you may be doing very complex joins across different tables.

    Andy Pavlo14:51

    I don't know yet whether agents are at that level or whether they're kind of going at it, really trying to read data against the database in more simple terms. So your question is, are agents going to change things? I think yes, right now, definitely in volume. And I think that's going to increase even further. And of course, now when the agents are running, they're generating their own telemetry, so that you have to store that and make sense of that.

    Agent memory: files vs. databases ("everything is a database")

    15:10
    Andy Pavlo15:10

    And that's another challenge. But in terms of whether the actual patterns of what they're doing, how different is that yet? That remains to be seen. This is something I'm trying to research as well.

    Matt Turck15:31

    Yeah, some fantastic stuff in there that I want to go back to in a minute. But there is a concept of agent memory that seems to be a little bit of a battleground right now, with, on the one hand, file systems, and on the other hand, databases. What is your view?

    Andy Pavlo15:50

    What is a file system? It's a database, right? Everything's a database, right? You have a directory of a bunch of Excel files. That's a database, right? You have a notebook, a pencil and paper. That's a database. When my daughter was like three, we were trying to teach her the importance of databases. And I would make her write down—well, we'd do it together—we'd write down the temperature every day in a little lined notebook.

    Andy Pavlo15:57

    And she'd close it and go, "Commit," to save the data. That's a database, right?

    Matt Turck15:58

    Yeah.

    Andy Pavlo16:09

    It really comes down to what the form of the input or the context you want to provide to your agents is. And as far as I know, it's all human-readable text, right? So whether or not that human-readable text comes from a bunch of JSON files you have in your directory, or it's extracted from a database system that can then convert that into the correct endpoint form, at the end of the day, it doesn't matter.

    Andy Pavlo16:49

    I think there's obviously advantages and safety guarantees you can get when you use a database management system that can protect data. SQLite is famous for this. Why would you want to use a file system and try to manage whether you're writing things correctly and safely, and that you can make multiple changes and commit this, versus just SQLite? SQLite just does it, right? So at the end of the day, everybody should be putting everything in a database system, right? Even some file systems, like WinFS, underneath the covers, it's a database system, right?

    Matt Turck17:23

    Yeah, there was a piece that also went viral at some point. I think that was entitled "I Recreated a Database by Accident," like talking about file systems. Okay, interesting. Does the world of multi-agents, sub-agents, agents, agent swarms, does that require more of a database? Meaning, they seem computationally very complex, and therefore having something that's extremely heavy-duty on the other side seems to make sense.

    Andy Pavlo17:52

    At the end of the day, if you need to share state across multiple machines or instances of these agents, whether they're running on the same box or not, again, a database management system provides that capability that you could have multiple readers reading the shared state, right? I haven't seen anything where the agent, as part of its memory lookup, is pushing down filters and things like that into the database. Usually that's provided as part of the context. Again, this goes back to the RAG or the vector database stuff. You provide the context based on what you've already looked up against your memory in the database files or the database itself.

    Which database do AI models recommend? The new SEO

    18:24
    Andy Pavlo18:24

    So again, from my understanding, the API for the input of the context is just human-readable text. And whether that comes from a file system or a database system, it doesn't matter. But the database system provides certain guarantees and luxuries, if you will, that a file system can't provide. And shared state is one of them across multiple agents.

    Matt Turck18:30

    Have you seen the agent world converge towards a specific type of database?

    Andy Pavlo18:38

    I mean, you're asking in terms of using it for memory, or having an agent choose one particular database system over another when they build an application?

    Matt Turck18:40

    More the—actually both.

    Andy Pavlo18:56

    So maybe the latter. The models are trained on what people are writing about. And so, if I say I'm building a new application, chances are it's probably going to recommend Postgres at this point.

    Matt Turck18:59

    Because that's in the pre-training data.

    Andy Pavlo19:20

    Yeah. And I will say, I can't say who, I've had two database companies ask me how to get their database to be recommended first by an agent over Postgres. And I was like, first of all, that's not what I do. That's not research. And two, don't talk to me. Go talk to the guys building the models because that's clearly something that they could have sway in. It's basically like SEO. How do you get the agent to learn, like, hey, I should use my database X instead of Postgres?

    "It cites me back to myself"

    19:25
    Matt Turck19:33

    I would suspect that your writing is highly indexed by the search engines, and therefore pitching you is a way of pitching the model.

    Andy Pavlo19:45

    So this is very much a first-world problem, and I'm not trying to humblebrag, but one problem I do have is when I ask questions about databases, it cites me back to myself because all my course materials—

    Matt Turck19:46

    Which has to be somewhat satisfying.

    Andy Pavlo20:03

    Yes. Well, no, because, say, when I was teaching database courses, I would say, "Show me—I want to learn about data systems that do this version of an index, like a B+ tree or whatever." And I would say, "So now I have to say, I already know which ones do do this, so show me the ones that don't do this." And then I have to say also, "Don't cite me. Don't cite Andy Pavlo or anyone from Carnegie Mellon."

    Andy Pavlo20:10

    But then the problem is that our course—

    Matt Turck20:13

    Does the model reply, "I know you're Andy," right?

    Andy Pavlo20:35

    ChatGPT, no. Yeah, ChatGPT actually says, like, "I know who you are." It says—I've said something like that. But Perplexity doesn't do this. But the problem is also, since all our course material is open source, anyone can just take and use it. So then it says, "Okay, well, you can't cite Carnegie Mellon University. Let me cite another university." But it's my slides being used at another course. So I gotta be very specific about what not to return back to me.

    Andy Pavlo20:59

    So I think that, to your original question, for OLTP, I haven't tried this in a while, but usually it spits back, "Hey, we should use Postgres." Unless you say something very specific, like, "I'm going to run on embedded devices," then it'll say use DuckDB or SQLite or something like that. Or, "I only care about key values," and it'll come back and say, like, use RocksDB. But for OLTP, it almost always says use Postgres.

    Andy Pavlo21:24

    And then it might recommend, like, "Do you want a hosted version or not?" if you have further chats. For OLAP, I'd be interested to see if there's any kind of ranking on this. It's actually not a bad idea. Maybe somebody's already done this. That one, it definitely would come back, and maybe you have to refine it further and say, like, "My workload looks like this, my data looks like this." And then I think it can give a more nuanced answer.

    Andy Pavlo21:45

    I haven't tried this for about a year or so, but I've definitely seen before it'll say, "Oh, do you have any money? Yes? Okay, you could use Snowflake or Databricks. Do you have no money? Okay, you could use DuckDB." Or, "Do you want to run on-prem? You could use ClickHouse," things like that. So then your second question was about agents?

    Matt Turck22:17

    Well, which one should they recommend? So that's the answer to which one they do recommend, but which one should it recommend? And where I'm going with this partly is this question of speed. And that's where ClickHouse shines. And we'll talk about ClickHouse more later. There's perhaps a concept of context layer ontology, which seems to be pushing towards graph databases. What makes sense?

    Andy Pavlo22:44

    I mean, I have strong opinions about graph databases. I don't think you need one. I have not seen any benchmarks that show native graph databases outperform relational database systems when those relational systems have the right constructs in SQL or the API so that you can do graph traversal all on the server side and not go back and forth. The SQL standard in 2023 added support for property graph queries. Oracle supports it.

    Andy Pavlo23:10

    Postgres was supposed to support it this year. It's coming out in Postgres 19, but I think it got delayed or rolled back. I don't know why. But anyway, that's beside the point. No, I think that the context layer, again, depending on the requirements of how many agents you have, whether you have shared state or not across multiple agents, Postgres often should be the first choice for many of these things. For OLTP. If you're doing analytics, that's a whole other story.

    MCP for databases, and text-to-SQL from 60% to 99.5%

    23:14
    Matt Turck23:16

    Let's talk about MCP.

    Andy Pavlo23:17

    Yes.

    Matt Turck23:44

    So over the last year or so, it seems like everybody, every database provider, added an MCP server to their offering. Thoughts on that, like pros and cons? And I'm coming from the year or two ago, there was this whole text-to-SQL kind of movement where everybody was trying to make LLMs work on top of databases, especially for analytical purposes, and that sort of didn't work. Are we there now?

    Andy Pavlo24:13

    So text-to-SQL, it's like natural language to SQL. So this is something people have been trying to do since the '70s, right? And every so often there's maybe some kind of breakthrough in natural language processing. People say, okay, now finally it can happen, right? It hasn't happened except for LLMs. I've been skeptical of the academic literature and the recent publications on various frameworks that do text-to-SQL because oftentimes the benchmarks are showing SPIDER and a couple other ones. Like, these are well-known benchmarks that the models are trained on.

    Andy Pavlo24:36

    So, of course it's going to regurgitate the correct answer. I actually was talking to somebody from a very large bank yesterday who said they now have a text-to-SQL interface. But in the initial implementation of it, they were using something off the shelf, and then they further refined it a little bit further in-house. But it was about 60% accuracy, meaning the LLM or the agent could produce the right SQL query, produce the right answer 60% of the time.

    Andy Pavlo25:06

    And that's roughly what I think I've seen before in some other reporting. 99.5%, which is insane. I think that's a viable thing now if you invest the time to have the right semantic layer and the context information you need, so the agent has enough information on how to derive the right SQL query from the question you're asking. I think, again, that's a large bank with infinite money, and they can pull this off.

    Andy Pavlo25:32

    I think if you just download something now, a random person will find that these existing tools are not going to be at that level of proficiency and accuracy. But it doesn't mean we can't get there, right? So now in terms of MCP, the Model Context Protocol, I get why they did it. I think this came out of Anthropic, and then everyone sort of quickly adopted it as the standard. And every so often there's a Hacker News post saying MCP's dead or whatever.

    Andy Pavlo25:57

    I don't think it's going away. It's basically a REST interface, right? A standard way to specify what you want the request to be from the agent to whatever data system you have. And then that data system that receives the MCP request, it's their job to convert it into whatever the correct API calls are to get the data you need. And in the case of database systems, it's going and writing the SQL query to get the data you need.

    Andy Pavlo26:25

    But usually it's oftentimes a strict translation of what's being passed to you and not trying to derive the semantic meaning of the MCP request, because the MCP request has to be somewhat well-defined. I have no problem with MCP. I think it's not this magic bullet that solves all the world's problems. It's just a digital interface, right, that you can now use to interoperate these different agents and different systems to collect the data you need.

    Andy Pavlo26:55

    So I'm all for it because you're using the standard rather than some homegrown thing. So yeah, fantastic. It's great. Could it be improved? I mean, I think it's inefficient because you're sending human-readable text over the network to go retrieve data. There's obviously better ways to do this with binary protocols, but at the end of the day, it's good enough, right? And it seems to be the standard everyone's using.

    Trust an agent the way you'd trust a junior developer

    27:09
    Matt Turck27:14

    Should it be that security layer that you were alluding to earlier?

    Andy Pavlo27:35

    Yeah. But it's no different than if the request came from SQL, from a traditional application client using JDBC or ODBC, the network protocols or terminals for these different data systems. It's no different if it comes from that versus a REST request with MCP. At the end of the day, okay, I got those requests. I need to understand who's asking for it, what are they asking to do, and should they be allowed to do it?

    Can AI build an entire database? Opus 4 and the CMU projects

    27:47
    Andy Pavlo27:48

    The guardrails we were talking about before, they're the same if it's coming in as SQL versus coming in as MCP.

    Matt Turck28:11

    Okay, so that's one part of the conversation. That's databases as memory layer or execution layer for agents. Let's flip the discussion into what AI can do for databases. So clearly, we live in a world where AI can code just about anything. Is that true of databases as well?

    Andy Pavlo28:30

    Yes. So, to give one anecdote, I teach a course on database management systems at Carnegie Mellon University. A year ago, the agents couldn't do our entire projects. So the projects would be like, we give you a scaffolding of a database system, and you have to implement the indexes, the query engine, and things like that. It could do some of it, not all of it. I think it was Opus 4, whatever Anthropic put out last year, that just opened the floodgates, and the agent could basically do all our assignments.

    Andy Pavlo29:01

    Right? With very little prompting. And of course, there is a lot of training data for it because all our projects are open source. They're all on GitHub, not just for students at Carnegie Mellon University, but also students outside of the university. We let them use it. So there's a lot of training data for them to implement things. And so I think agents basically could implement anything you would want to build in a database system now. Obviously, you have to prompt it the right way and hold its hand and make sure you generate the right design or it produces the implementation based on the design that you want.

    Andy Pavlo29:15

    At the end of the day, I think these agents are very capable of being able to do this.

    Matt Turck29:35

    So they could build an entire database. Because my experience of building databases as a venture investor is that it's a 10-year journey of pain where nothing much happens for three years, and you have some of the smartest people in the world getting together to solve what seem each time like insurmountable problems. So we're now saying that it can build the whole thing. Yeah.

    Andy Pavlo29:55

    The old adage from database systems is that it takes 10 years, but it's a system. You can build the first 90% in three years, and then the remaining 10% takes the next seven years. So yeah, I think that the agents are very capable. With enough tokens, of course, and with enough guidance, you can vibe-code an entire data-driven system. And there's certainly companies that are doing this now, and pretty much every single data-driven company is using agents to help develop things.

    60% of open-source databases now have AI commits

    29:57
    Andy Pavlo30:22

    And one of the things we do now is we keep track of every single database system that I know about. And for the open-source ones, every single night we pull down all the latest commits on GitHub, and we then track to see which ones are actually being co-signed by Claude or Codex and things like that. And obviously, if people turn that feature off, you don't know whether it's actually being generated by an agent, but most people don't do that.

    Andy Pavlo30:38

    And at this point, I think over 60% of the open-source database systems have commits coming from agents.

    Ten years of self-driving databases, Peloton to today

    30:43
    Matt Turck30:58

    And so does that, beyond the writing, also apply to the running of it? So, going to that self-driving database concept, that's something that was a big project of yours 10 years ago, I believe.

    Andy Pavlo31:00

    Roughly, yeah. Peloton.

    Matt Turck31:07

    Yes. So walk us through that journey. What was not possible then that has become possible today?

    Andy Pavlo31:30

    Yeah. So when I started at Carnegie Mellon University, one of the things I did in my first years was go visit companies and sort of see what sort of challenges they were facing with databases. And the overarching theme I saw over and over again was just running these systems, maintaining them, and optimizing them was a huge struggle. And this is not a huge revelation for me. People have been trying to do this for decades. I mean, since the creation of the relational model and relational databases in the 1970s, people have been trying to do autotuning for indexes, partitioning keys, sharding keys, and tuning knobs and so forth.

    Andy Pavlo31:50

    Microsoft Research did a lot of work in the early 2000s in this AutoAdmin project. They had a bunch of tools that allowed you to manage and optimize databases automatically.

    Matt Turck31:54

    And for context, that is because there's thousands and thousands of possible configurations of a database.

    Andy Pavlo32:17

    Right. So there's, like, what indexes should I have on my table to speed up queries? And obviously, if you have too many indexes, the writes go slower and you run out of disk space. There's the various knobs that control the runtime behavior of the system, how much memory to use for one memory pool versus for indexes and so forth. This is beyond the capabilities of any one human to reason about. And then if you start looking at if you have thousands of databases, it's just not possible for anybody to manage these things.

    Andy Pavlo32:50

    So we spent some time looking at how to build automated tools using machine learning and then what you may be able to call AI now, but using automated methods, machine learning techniques to automate this. And we sort of had two research tracks. We had one where we sort of took an opaque-box view of the database system where we said, we can't change the internals. What can we do just with the APIs that the system exposed? So, basically, how do you optimize automatically existing databases, MySQL, Postgres, Oracle, and so forth?

    Andy Pavlo33:20

    And then another track of research was, if you build the database from scratch, assuming it was going to be controlled by machine learning tools, or now we'd call agents, how would you design the system slightly differently? And a lot of things we learned from trying to optimize MySQL and Postgres, we ended up building in our own system. So the challenge always was then for these automation tools just having enough training data of production databases and being able to take the things you learn optimizing maybe one deployment and apply it to another.

    Andy Pavlo33:51

    Like, as I was saying before, most people have a staging or development environment, and then they have the production database. And so you obviously don't want to try random things on the production database that may cause you to slow down and degrade performance while you're trying to figure out how to tune it. So you would vet things on the production—sorry, the staging database—and then try to apply it to the production one. But oftentimes, it's never the same hardware, never the same workload.

    What LLMs changed: 85% of the tuning in 15 minutes

    33:52
    Andy Pavlo34:18

    So that was a big challenge we were always facing, just getting enough training data to feed into the models that can then make good decisions on the production databases. And then the LLMs show up, and that really changed how you can approach the problem because they are basically trained on blog articles, documentation, best practice guides that are all available on the internet of how to tune these databases. And so, on the research side, what we found is the LLMs can get you maybe 85% of the way there in tuning certain things.

    Andy Pavlo34:45

    And obviously, there are weird corner cases; bespoke algorithms can handle those. But our models oftentimes would take hours and hours to train, assuming you had enough training data to produce the pristine optimal configuration. But the LLMs can come in, just like in 15 minutes, and produce something that was good enough. And that's good enough for most people.

    Matt Turck34:53

    So are we there yet? Are we in a world where databases can largely be sort of self-tuned, self-optimized?

    Andy Pavlo35:04

    I mean, the answer is yes, but of course it depends on how much improvement you actually want. Like I'm saying, the LLMs can get you 80, 85% of the way there. Is there always more performance to squeak out?

    Matt Turck35:05

    Yes.

    Andy Pavlo35:18

    And some people are always going to care about these things, and therefore the automated methods we built in the past would matter a lot more. But for most people, actually, one of the things I was surprised so much was when we had the startup and were trying to do this auto-tuning stuff for PostgreSQL, is how many—

    Matt Turck35:19

    That was OtterTune.

    Andy Pavlo35:36

    OtterTune, yes. But how many people were running with the out-of-the-box configuration from Amazon or Google in their cloud platforms? And they would tell us, "Oh, we thought Amazon was tuning this for us." And no, they're not, right? So there's enough people that are maybe running with the default configuration, which is, that thing is just terrible. Don't do that. An LLM could come and take care of some low-hanging fruit and provide some meaningful improvements for people in a short amount of time.

    Should anyone still study databases?

    35:42
    Matt Turck36:13

    So, in the rest of the software world, we've seen the nature of the job of a developer profoundly change. If you're a student at CMU today and you're super excited about the world of databases, so yes, you can work for the ClickHouses or CockroachDBs of the world, but outside of that, what is a job today that requires deep knowledge of databases if we're in a world where agents are going to be, like, calling them, managing them, and then the databases themselves will be self-driving?

    Andy Pavlo36:30

    So, what I would tell my students always was, for my courses, even if you're not going to go off and become a database system engineer, like go work on the internals of ClickHouse or CockroachDB or Postgres, at the end of the day, everyone's going to interact with a database system. And if you understand the internals of them, how they're implemented, what the data system is trying to do or not do for you for given queries, you're just going to be in a better position to understand what's going on in the world in tech stacks.

    Andy Pavlo36:53

    Now, do you need my course to figure that out? No. I'd be naive to think so, that the agents, they've sucked all my stuff in so they can regurgitate everything I would teach anyway. So, but at the end of the day, again, whether you're learning it from a traditional university course or self-taught or having an LLM teach you things, if you understand what a database system is trying to do, then it just sort of removes months of mysteries and helps you understand what's going on in your application. Why are queries running slow or not running slow?

    Andy Pavlo37:32

    Just, you're in a better position to understand what's going on. And so, if you pursued a computer science degree or engineering degree and focused on database management systems, I don't think the era of, like, you're gonna write a system from scratch by yourself, I think that—I'd be naive to think that that's still going to be around. But at the end of the day, the database companies are still hiring people to work on these things because you need someone to understand what the agents are actually doing and make sense of it.

    Matt Turck37:53

    But to fully answer the question, like, we're in an era where there's going to be just no humans in the loop in the way databases are managed and deployed.

    Andy Pavlo38:10

    I would say that the traditional role of a DBA, as we've known in the past, I think that's going to be relegated. And a lot of that can be automated. Doesn't necessarily mean those roles go away. They just now get to pursue more higher-minded activities. Like data modeling is always, always important. And again, agents can do a lot of that too. But you sort of need someone to understand, like, yes, does this data model capture what our organization or business is trying to represent in the database?

    Free CMU courses, the DJ, and the Wu-Tang final exam

    38:27
    Andy Pavlo38:28

    All those things are super important still. Just like, does someone need to be tuning the knobs or picking exact indexes for tables? I think that ships out. Agents can do all that.

    Matt Turck38:55

    Great. To switch topics, a lot of people in the database world, I assume pretty much anyone, is very familiar with your work. But for people who are not, you have been a tenured professor at CMU. You have had very popular classes. I think part of the one beautiful thing that you did is that you took a lot of CMU courses that you were teaching and then put them online for free?

    Andy Pavlo39:13

    So when I started at Carnegie Mellon, I was the only data systems professor there. There was another professor there, but he did more graph mining, data mining stuff. And so when I started in 2013, I was like, okay, well, I got to figure out how to get tenure. And competition is not the right word, but I'm competing against MIT, Stanford, and Berkeley, and I'm the only data person there. So, all right, well, let's just put everything for free on the internet.

    Andy Pavlo39:36

    And help promote our research, what we're trying to do. It also had the side effect of, like, I would end up being more prepared for classes because I would—not that I was phoning it in anyway—but I would make sure I had my stuff together because I knew it was going to be recorded and watched, because I don't want to say the wrong thing. I would be more professional. I would say less crazy things, which is, people are shocked when I say this, that my course even now is considered watered down.

    Andy Pavlo40:03

    Like, even for them, that's bizarre. So yeah, obviously the benefit is not everyone can go to Carnegie Mellon University. It's not cheap. It's a private school. Not everyone can get in, but why should that education be held back from other people? So it doesn't cost me anything. I just put it on YouTube and make everything available to everyone else.

    Matt Turck40:18

    And we were joking about Wu-Tang Clan at the beginning of this conversation. I think I heard or read somewhere that you had a DJ spinning Wu-Tang Clan at the beginning of class. What is that story?

    Andy Pavlo40:35

    So it started in 2019. I was like, all right, well, let's have fun with this. And when you set up the cameras and things like that for your lectures, there's obviously time to set things up, the wires—we had to set all this up too, right? There's time. So I'm like, all right, well, this is awkward silence where the students are coming in, they're staring at you. There's nothing really going on.

    Andy Pavlo40:53

    There's nobody saying anything. So, like, let me get a DJ to sit next to me and just play beats and play music next to me the whole time, right? And then you start talking to them a lot during the lecture and you find out they live pretty interesting lives. Like, one guy had a bunch of ex-girlfriends trying to hook up with him. One guy would loan a bunch of money to people he shouldn't be loaning to.

    Andy Pavlo40:59

    So we just talk about it in the class, plus databases.

    Matt Turck40:59

    As one does.

    Andy Pavlo41:00

    As one does.

    Matt Turck41:00

    Yes.

    Andy Pavlo41:14

    And then I think one year we had—I did a little joke where oftentimes the top database conferences are always at the beginning of the semester, at the end of the semester. So I have to travel. So the first week of classes, I'll be traveling to go to a conference. So you don't want to cancel class because it's super important for students to know what the course is going to be about the first week because they're trying to figure out classes they want to take.

    Andy Pavlo41:39

    So I would film it remotely. So one year I was in L.A. for a conference. And then at the very end of the lecture, because I filmed it, I was like, oh, by the way, the most important things you need to know about databases. And I just listed all the members of the Wu-Tang Clan on the first album, plus Cappadonna, who was in jail at the time. And I said, this is the most important thing you have to know about databases.

    Andy Pavlo41:57

    And then when it came to the end of the semester, the final exam, the very first question was: list all the members of the Wu-Tang Clan on 36 Chambers, plus who was in jail. And you have to get exact spelling. And if they got that question right, they would have got 100% correct on the exam.

    Matt Turck41:58

    Yeah, not killer, but killa.

    Andy Pavlo42:08

    That's right. So someone got that wrong. They missed the H when they should have had it. So no one actually has ever got that. And obviously I can't do it again because students are expecting it.

    Why start a research lab inside ClickHouse?

    42:09
    Matt Turck42:22

    Okay, amazing. All right, so that takes us in some way to ClickHouse and ClickHouse Labs. You joined ClickHouse recently to start ClickHouse Labs.

    Andy Pavlo42:22

    Yes.

    Matt Turck42:28

    So why start a lab in a commercial company versus doing that in academia?

    Andy Pavlo42:45

    I mean, without going too much into politics, the nature of research in universities is kind of tough right now. The funding in the United States is not what it used to be. There's also a larger question of, like, what does it mean now to have a PhD student when an agent can kind of do a lot of things, right? Do you need to have as many PhD students working on stuff when an agent can produce things and work with me as the professor?

    Andy Pavlo43:14

    To write papers and things like that. I mean, there's certainly still value in educating students, and I like doing that, but it just wasn't clear to me, like, how would I raise money to have as many students as I would want? So I had some companies reach out to me about, like, hey, are you interested in going and doing some interesting things with us? But it was always a bit more like, hey, do you want to come be an engineer with us, or kind of watered-down versions of what I wanted to do on the research side.

    Andy Pavlo43:47

    And then when I talked to the ClickHouse people, it almost seemed too good to be true because they were like, hey, come do research with us. I'm like, okay, this is what research is. You understand this? They're like, yeah, yeah, go do this. I'm like, okay, do you understand, like, things might not always pan out correctly because this is research? I don't know. Like, yeah, great. Do it. Do this. And it's insane because they're not an early-stage startup, but they're a startup.

    Andy Pavlo44:02

    They're not a public company. And normally you only see companies, at least in the database world, set up research labs when they already get established.

    Matt Turck44:03

    Right.

    Andy Pavlo44:21

    And for them to take this kind of risk, or to have the luxury of having someone like me come in to do pure research, is amazing, and they're very serious about it. So the opportunity almost seemed too good to be true. And so far, I'm two months into this, and it seems to be working out. I'm sort of in this observation period right now where I'm going through and understanding the stuff that they've built so far because they haven't really published it or written about what they've done.

    Andy Pavlo44:53

    They have some blog articles which are super interesting, but they've done some amazing work that I certainly can't take credit for doing. My job right now is sort of helping them turn those into publications we can disseminate and put out into the world, as well as start ramping up longer-term, more exploratory things. To be very clear, even when I was at the university, in database systems, my work is very much applied.

    Andy Pavlo45:18

    So it's not like I'm doing theory work that only 10 people in the world can understand. It's not directly related to anything that's going on. The kind of stuff I want to do at ClickHouse is directly relevant to what ClickHouse wants to do, not just on the ClickHouse system itself, but also now they have a Postgres offering. There's a bunch of stuff we've done in the past on Postgres, and now we can start looking at real data, trying to put some of these things in production workloads and help people put things out there.

    Why ClickHouse looked like vaporware in 2016

    45:53
    Andy Pavlo45:53

    So that, to me, is super exciting: to have a direct impact. Instead of me writing the paper and then hoping someone comes along and takes it and adopts what we're doing, we can do the research and evaluate whether it's a good idea or not on real workloads, on real production systems. And if it does make sense, then we can push it out in a short amount of time. So that part is super seductive to me, that I have this opportunity.

    Matt Turck46:13

    Okay, fantastic. While we're at it, let's talk about ClickHouse a little bit, in which, for disclosure, we are very proud investors here at FirstMark. I think I read somewhere that in 2016 you were concerned that ClickHouse was possibly vaporware because it was a little too good to be true.

    Andy Pavlo46:13

    Yeah.

    Matt Turck46:15

    Can you unpack that for us?

    Andy Pavlo46:38

    Yeah. So, as I mentioned, my job as a database researcher is trying to understand what else is out there, because databases are a really interesting research area. Not only am I, again, using the term "competing" loosely—I don't mean that in a pejorative sense—but you are in competition with other researchers trying to get ideas out and published. So not only am I competing with other researchers at other universities,

    Andy Pavlo47:05

    you're also competing against the big tech companies—the Microsofts, the Googles, and Amazons of the world. Plus, there's all these startups that have their own database systems. There's not a lot of operating system startups. There's a lot of database system startups, as you know, as an investor. So my job is to try to figure out, when a new system gets announced, what are they doing differently than my own research? What are they doing better?

    Andy Pavlo47:11

    What can I learn from how they approach the problems that they're trying to solve? And so, when the announcement of—or the release, the unveiling of—ClickHouse in 2016 came out, I was like, I know how hard it is to build a data system, and for this thing to appear out of nowhere and have all these capabilities, like the vectorized execution, the columnar stuff, the compression things, the compaction stuff, it almost seemed too good to be true, or it's a fork of something that already exists.

    What makes ClickHouse fast: columns, vectors, Snowflake's lineage

    47:46
    Andy Pavlo47:47

    But at the time in 2016, there wasn't an OLAP engine that was open source that would have all these capabilities, right? So that threw me for a loop too. I was like, oh, so I just assumed it was fake. And then sure enough, it's real. Obviously, it has a lot of benefit.

    Matt Turck48:03

    So what are some of those benefits in particular? It's very fast, right? It's very real-time. Why is that possible? Why do databases typically struggle with that, especially analytical databases? And why are they able to do that?

    Andy Pavlo48:30

    So I would say the architecture of ClickHouse, the core fundamentals, by 2026 standards, is not novel. By 2016 standards, certainly, it was. The nature of, like, you can store data in this columnar layout. So it's storing all the attributes, all the data for a single column together contiguously, versus, like, in a row store, you store all the values for all the columns of a single tuple contiguously. Like, storing things in that manner, using vectorized instructions in CPUs to process multiple pieces of data at the same time within one single instruction, versus, like, having to do for loops and run things less efficiently.

    Andy Pavlo49:04

    The built-in compression mechanisms, the way they do a log-structured storage where you're appending a bunch of changes and then, because you want to ingest as fast as possible, eventually in the background you'll compact it and store it in a more efficient manner. A lot of those ideas of the core architecture come from the work done at Snowflake. And prior to Snowflake, there was another system called VectorWise, or X100—sorry, the MonetDB/X100, which then became commercialized as VectorWise from CWI, which is the same school where DuckDB was built.

    Andy Pavlo49:42

    So a lot of the ideas that people have been developing for these column-store systems and analytical systems and real-time analytical systems had been sort of floating around, but nothing was open source in 2016, at least as far as I know, that did all these things. So when ClickHouse came out of the gate like, hey, we have all this stuff based on all this research that people have done before, it was amazing, right? And since then, DuckDB has something similar, Polars from the guys in Amsterdam, Firebolt started off as a fork of ClickHouse and they rewrote a bunch of stuff inside, and now they have all these things. Like, Databricks has their own version of the engine.

    Andy Pavlo50:23

    So everyone has this sort of architecture that is predicated or based upon the work done from Snowflake and VectorWise. But again, in 2016, that was super novel. Now it's sort of table stakes. And one of the things I think ClickHouse has done really well over the years is just keep improving the performance even further and expanding what workloads they can actually support.

    Matt Turck50:28

    And what use cases are you seeing for ClickHouse from the inside? What do you see customers do?

    Andy Pavlo50:47

    So I think going back to the agent stuff, the agent stuff seems really interesting too, that people are using this to observe the information, the telemetry that these agents are generating, and be able to react to them in real time and ask questions about them. There's a bunch of stuff done in the financial markets as well, fraud detection. Again, anything where you want to be able to ingest data very quickly and then ask questions about it, complex questions about it, very, very quickly and get back things in milliseconds rather than seconds.

    Andy Pavlo51:03

    Like that. I think the combination of doing those two things puts ClickHouse, I think, in a unique position.

    Matt Turck51:27

    And ClickHouse has been on a journey to expand its product offering, right? So it started with this very powerful OLAP analytical real-time database, and now it does—it bought Langfuse for LLM observability, bought other things. Like, how do you from the inside understand this evolution? I mean, does it all fit together?

    Andy Pavlo51:50

    I mean, somebody asked me, like, why did I join ClickHouse? Again, I'm only two months into this, and it's kind of like asking, like, you just got married and somebody asked you, why'd you choose your wife or husband? I'm like, oh, because they chose me, right? Like, it's a bit more than that. No, honestly, before I decided to join ClickHouse, they were doing things at the business level. Like, so this is ClickHouse Incorporated, not ClickHouse the system.

    Andy Pavlo52:20

    They were doing things at the business level that I would've done if I was running a database startup like they were, like having a Postgres offering, not just a single OLAP engine that's really good. Now starting in the OLTP space, having the monitoring, the visualization, the observability stack, not just for traditional BI or traditional data, but also now the agents and the LLMs. Like, they're doing all the things to expand out different verticals that build on their platform. Like, all that, to me, that all makes sense and tracks.

    Why Databricks, Snowflake and ClickHouse all added Postgres

    52:22
    Matt Turck52:50

    Maybe segueing into the state of the database market in 2026. We've alluded to a bunch of this already, but I'd like to put it all together. So one thing that ClickHouse has done, as you mentioned, was adding a Postgres service, a managed Postgres service. It seems that it's a clear trend. So Databricks bought Neon. I think Snowflake made an acquisition.

    Andy Pavlo52:51

    Crunchy Data last year.

    Matt Turck53:08

    That convergence of analytical databases and transactional databases, is that the future? Is that driven by commercial and business reasons so you can sell more to the same customer, or is there an industrial sort of product logic to it?

    Andy Pavlo53:22

    There's two things. Should a database company offer both an OLTP transactional operational data system and an OLAP system, or at least whether or not they're separate systems or not, that's up for larger discussion. Should you do that?

    Matt Turck53:23

    Yes.

    Andy Pavlo53:50

    I mean, going back to what agents allow you to do, obviously maintaining these things is still a challenge, but the bar of entry of deploying a Postgres offering has certainly been reduced. Another conjecture I have in my mind is, is it better for an analytical company, a startup that starts with an analytical database, to then later add an OLTP offering versus an OLTP company then later adding an OLAP offering? Again, I realize I'm saying this and I'm biased because I'm at ClickHouse, but you look at what Databricks did, you look at what ClickHouse has done.

    Andy Pavlo54:22

    To me, that appears to be the right path to do this. If you want to start a new data startup, start doing analytics first, get traction there, then you can add the OLTP offering. But I can't prove why. Just from the business perspective, it seems to be the smarter play. Now the question is, okay, should you have a single API, single interface, or a single logical endpoint that allows people to run transactions and analytics on what is perceived as a single database instance, or should you have separate services?

    Andy Pavlo54:49

    ClickHouse right now is separate services. But they have a fast path to get the data out of the OLTP side from Postgres into ClickHouse. What Databricks announced this year was their LTAP engine, or LTAP story, for their Lakebase, where you can run your transactions on Postgres or Lakebase or Neon, and then have the OLAP engine be able to read that data directly so that you don't have to make two copies of it in the way ClickHouse does or other systems do.

    Andy Pavlo55:33

    So that idea is not new, to have this sort of two engines but one copy of the data. There was a previous startup in the 2010s called Splice Machine, which actually I was an advisor for because the founder was a CMU alum, and he asked me to be an advisor for them. But they were running HBase plus Spark SQL, and they had a single copy of the data that was a row store. I mean, people have been wanting these HTAP systems, or these hybrid systems, since the very beginning.

    Andy Pavlo56:07

    It's just the challenge has always been, how do you reconcile the fact that some of the data might still be in the row store versus the column store? And the major commercial vendors like Oracle and SQL Server, they'll store two copies, one in the row store, one in the column store. It's called a fractured mirror approach, and they have an internal mechanism to synchronize these things. But the challenge has always been why this sort of hybrid approach has never taken off is because the stakeholders at organizations or companies for these two different categories of workloads have always wanted their own best thing.

    Andy Pavlo56:48

    The OLTP team, the people that are running the application, the operational side, they don't want something that kind of does okay at operations and can do okay at analytics. They want the best operational database system. And for better or worse, that's Postgres now, right? And best is not always in terms of performance. It could be a variety of reasons why it's considered the best. And same thing for analytics. You don't want a system that kind of does okay at analytics and does okay on operational workloads.

    Andy Pavlo57:23

    I want the best analytical system. And so that's always been the business challenge, the go-to-market challenge for these hybrid systems. I think, though, Databricks might be in a good position to potentially pull this off because they already have that sort of market from the analytical side, machine learning side, and the data warehouse side. And now you're just adding in this OLTP side. And so ClickHouse could potentially pursue the same thing. That's getting kind of in the weeds of database system architectures.

    Andy Pavlo57:39

    But I think everybody's always wanted this. The challenge, though, often isn't always just pure engineering. There might be organizational reasons or human reasons why you can't achieve this.

    Matt Turck57:51

    Yeah, super interesting. So to play it back, you could abstract away the complexity to the user, meaning that you could have presumably an LLM that translates your query—

    Andy Pavlo57:51

    Like a router.

    Matt Turck58:03

    Yeah, exactly, a router. However, from a fundamental architecture perspective, it's very hard to combine. And from an organizational perspective, it may not make sense because people want the best of breed.

    Vector, graph and GPU databases: thumbs up or down?

    58:15
    Andy Pavlo58:15

    Yes, but Postgres is the best of breed now in many cases. So betting on your front end as a Postgres deployment that is compatible with Postgres, that might overcome that issue.

    Matt Turck58:24

    So, to talk about other parts of the database market, we talked about vector databases. Is that thumbs up or thumbs down?

    Andy Pavlo58:25

    What is thumbs up, thumbs down?

    Matt Turck58:34

    Like, in terms of, is that a good business to get invested in, or are they going to be around now that everybody else and their brother has a vector search capability?

    Andy Pavlo59:01

    One of the things that would happen before when I was a professor is, you always hear rumors about who's doing well, not doing well, through a combination of either the investors or former employees or students that maybe go to internships or whatever, or maybe interview some places and they come back. So you get sort of bits of information from everyone. You kind of piece together what the data landscape looks like. Problem is, now we're in ClickHouse, now I see everything.

    Andy Pavlo59:25

    As an investor, you see everything too. The one vector database company that I know is doing very well is Turbopuffer, and they are hyper-specialized in doing vector search at a cost-performance ratio that's much better than everyone else. So I don't think that the vector databases are going to go away. I think that they'll evolve in two ways. They'll have to become either a sort of general-purpose system like a Postgres, like a MySQL, where they become the system of record where you're storing the original tuples plus the embeddings or the vectors for them.

    Andy Pavlo1:00:04

    Or they become like an Elasticsearch, where there's a separate system where they have a copy of the data that's being pulled from the operational side. And in that case, they can live sort of comfortably as being this additional thing you add on. And if you want the raw best performance of vector search, in some cases you may have to go to one of these specialized systems. So I don't think that's going to go away. I just don't think I've seen predictions that, like, oh, Postgres is going to die at the hands of a vector database.

    Andy Pavlo1:00:13

    That's not happening. That's not happening.

    Matt Turck1:00:14

    Very much the opposite, right?

    Andy Pavlo1:00:15

    Yes.

    Matt Turck1:00:22

    Graph databases we mentioned at the beginning. So, not to pick on them, but Neo4j has been around for 20 years now.

    Andy Pavlo1:00:23

    Sure.

    Matt Turck1:00:28

    Yes. And this was supposed to be the moment for graph databases: AI. So what's happening there?

    Andy Pavlo1:00:52

    I have an outstanding bet with somebody on Hacker News where they said that by the year 2030, the graph database market was going to be larger than the relational database market. And if this becomes true, then I will wear a shirt that says, "I love graph databases," and I will use that as my driver's license, my university ID. I'll put it on my website till the day I die. I'm pretty comfortable. It's 2026.

    Andy Pavlo1:01:16

    We got four years to go. This is not happening. No, it's always been a niche market. And I think that, because my perspective on the research side, the research shows that if you do certain things in implementing the engine, which ClickHouse does do, DuckDB does some of this as well, there's things you can do that allow you to do the traversals of graphs, which are essentially just joins, self-joins on a table. You can implement those things very, very efficiently, and you can easily outperform Neo4j.

    Andy Pavlo1:01:42

    And that's kind of a kicking—like Neo4j, saying you're faster than Neo4j is like saying I'm faster than somebody that's in a wheelchair, right? You can run fast. It's a low blow. So, but I'm just saying that all the graph databases, I think even the best ones, you're just not gonna—

    Matt Turck1:01:42

    Yeah.

    Andy Pavlo1:01:58

    There's no system that has as much of these optimizations that are in the research and actually appearing in some of these systems now. You're gonna lose. What you will lose against a graph database is if you're doing the graph traversal with the client side and the server side, meaning, like, I got to figure out what the next node I want to go look at. I go back to the client and that decides the next node to go traverse.

    Andy Pavlo1:02:18

    If you're doing that back and forth, yeah, they'll beat you guys. But like I said, the SQL standard now supports property graph queries. Oracle has this, right? They were a big pusher of this. This extension of SQL allows you to do that traversal on the server side. So graph databases, I think, are a horrible idea. There's no reason why anyone would actually want to use them.

    Matt Turck1:02:20

    GPU databases?

    Andy Pavlo1:02:20

    Yes.

    Matt Turck1:02:31

    I know you have a special interest there. There was a cycle when the generation appeared, then went away. There seems to be a renewal. What is a GPU database, and what is your prediction?

    Andy Pavlo1:02:56

    So, a GPU database is a data management system where the execution engine for queries is offloaded to a GPU running on PCIe, or running in the same box or another box. So, the history of people trying to build accelerators for data systems goes back to the beginning of data systems. In the 1970s, they were called database machines. So people would build specialized hardware to run sorting and query execution operators. And that obviously died out in the early 1980s because by the time it'd take you to design and fab new specialized hardware, Intel or Motorola would put out the next CPU, or the hardware got better and just the gains you were getting went away.

    Andy Pavlo1:03:33

    So hardware accelerators for databases basically died out in the 1980s. There wasn't a lot of activity in the '90s, 2000s. You saw sort of the rise of people trying to do FPGAs for databases, and every so often that comes back now. Some of the cloud vendors do a little bit of these things, but usually to filter things on the NIC, on the network side of things coming in. So, with GPU databases, again, there was a bunch of systems in the 2010s that were trying this.

    Andy Pavlo1:04:04

    We did a seminar series at the university where we invited all the GPU database guys to come and give talks about what they were doing, why they were faster than existing systems. And the big challenge at the time was that those systems, you had to put the entire database inside the memory of the GPU. Because if you had to go back up through PCIe, it was just way too slow. And a bunch of those startups sort of fizzled out. Some of them are still around, but they're sort of specialized for doing visualizations.

    Andy Pavlo1:04:36

    And then there was, in the last year or so, NVIDIA has basically gobbled up a bunch of these GPU database companies that were kind of struggling along. And I was an advisor for one of them called Voltron, but they also picked up HeavyDB. And so NVIDIA is all in on this now. So it remains to be seen whether the idea that you're going to build a CPU-only database system, long-term, whether that's going to still hold. I've heard, again, mixed reports. This is public.

    Andy Pavlo1:05:01

    Microsoft has offerings now in the cloud that can be accelerated for your data system, can be accelerated with GPUs for analytics. Another major database company that I can't say who they are, they looked at the economics of GPUs and decided it wasn't worth it. So one of my former students, now a professor at the University of Wisconsin, they're now on leave at NVIDIA. They have a project called Cider, which is not necessarily a new database system, but it's a layer in between an existing system and CUDA.

    Andy Pavlo1:05:34

    And so it supports taking DuckDB queries and running that down on the GPU. I think they can do this in Doris or StarRocks and DataFusion. And so at ClickHouse, we've been potentially looking at this as well, but it's research. I don't know. It's interesting to see how much translation you have to do between how ClickHouse expects things and how CUDA wants things to be, the data layout and so forth. How do you organize memory or share memory between these different components?

    Andy Pavlo1:06:05

    TBD. It remains to be seen whether this actually makes sense. But certainly, there's a lot of research energy behind this. And publicly, I can say this: NVIDIA is obviously pushing this because it'll sell more GPUs, right? Because it's hard enough to get new CPUs. Everyone's compute-bound, or memory is hard to get. The computing hardware is very expensive, hard to get now. And GPUs, of all the things, are the most expensive hardware to get. And now you're going to say your entire database is going to run off GPU.

    Andy Pavlo1:06:20

    I don't know if that makes sense, at least in the short term. But if the performance improvements are quite significant, and some of the research shows that it is, maybe it makes sense.

    Matt Turck1:06:28

    Is there an emerging category, or maybe a niche somewhere within a category, that people don't talk about enough yet?

    Andy Pavlo1:06:29

    Of databases?

    Matt Turck1:06:30

    Yeah.

    Andy Pavlo1:06:51

    I mean, you can always build new data systems for new hardware, but my track record on this is terrible. We've done a bunch of research on experimental hardware, and it always gets canceled. Or it's even not that experimental. It's like Intel had this Optane persistent memory stuff. We did a bunch of research on building systems for that. Because if you assume now your DRAM is persistent, like you pull the plug and you don't lose anything, that changes how you fundamentally build a data system.

    Andy Pavlo1:07:24

    We did a bunch of work on that, and then Intel killed that product line. We were doing other research on processing-in-memory hardware. So think of DRAM sticks with CPU cores directly on the DIMM. So the data system now can say, okay, instead of pulling things from memory, bringing it into my CPU caches, and then I can compute things on them, I'll just send the query down to the DIMM itself and run it there. We were doing a bunch of work on this thing called UPMEM that got bought by Qualcomm and got killed last year.

    Andy Pavlo1:07:53

    So that didn't work out. So there's always a bunch of work you can do on data systems, on new hardware. I would say that, actually, I don't know the answer, right? This is one of the things I'm trying to figure out at ClickHouse. And we talked about this in the very beginning: are agentic workloads significantly different than what humans or what existing applications do now? And if so, why or how? And how would you change maybe the development of a database system to take better advantage of this?

    Andy Pavlo1:08:19

    That remains to be seen, how that works. I think there's always a bunch of problems in query optimization that I think are interesting. That remains the hardest part about database systems. Incremental materialized views, another big challenge. Again, these are not things that no one else has thought of. People have been trying to do these things for decades. So I think the agentic stuff is probably the most interesting and relevant thing to me right now: what changes with these workloads, and what changes in the system architecture?

    Is the database market stagnant? "A cheetah on cocaine in a Ferrari"

    1:08:34
    Andy Pavlo1:08:34

    And to be honest, I don't know the answer because I just haven't seen it yet.

    Matt Turck1:09:03

    I'm asking because it generally feels like the database market is at a specific moment in its history, meaning that there was an explosion of activity in the 2010s. There was SQL versus NoSQL, that whole evolution. Then there was the emergence of Databricks, Snowflake, now ClickHouse. But it seems that things have slowed down a little bit in terms of explosion, I guess.

    Andy Pavlo1:09:29

    Stagnant, maybe. Yeah, no, but we've been through this trend before, right? There was a lot of activity in relational databases in the 1970s, 1980s, and then the 1990s. Again, people sort of, the market sort of solidified around these major enterprises: the Oracles, the IBMs, Teradatas. And then if your only viewpoint of databases were from those kinds of companies, then yeah, it looked like it's been stagnant for years. But as you said, a lot of activity in the 2000s, 2010s.

    Andy Pavlo1:09:55

    I mean, relative to AI, then everything looks stagnant because there has, in the history of computer science, been nothing like that before. Like, it's just like taking a cheetah, giving him a bunch of cocaine, and putting him in a Ferrari. Like, the amount of speed that people are developing these things is insane. So everything looks slow or dead or stagnant to that. But at the end of the day, I think the volume is going to matter a lot.

    The relational model is arithmetic; SQL as the new assembly

    1:10:14
    Andy Pavlo1:10:26

    I think that one interesting question is, to your point of cost and efficiency, squeaking out the best performance you can get for the hardware that you have, trying to reduce that cost. That's always an interesting challenge that you could pursue. But at the end of the day, I don't think there's going to be a massive change in what data looks like that requires us to throw everything away that we've known about databases. In the same way, you wouldn't come up with a new notion of arithmetic or mathematics to replace one plus one equals two.

    Andy Pavlo1:10:57

    The relational model itself is the foundation of how you want to represent data. You can vary the implementations of that. And so there's certainly a lot of work to make these systems more efficient for this. But I don't think you're going to throw everything away and the agents need some kind of database system that you've never even conceived of before. At the end of the day, it doesn't make sense. So I don't think that part changes.

    Andy Pavlo1:11:20

    I think, like I said, there's always new hardware. I think there's certainly improvements that can be done for SQL. There's always going to be people trying to replace SQL. I think that might be a lost cause, although SQL could end up being like how, in the same way that people don't write assembly anymore, SQL might end up being like that because text-to-SQL works so well. I think the agentic stuff is super interesting, and I think getting better performance is always going to be, at least from my perspective, a fun challenge, just things we can pursue.

    Larry Ellison, Linux, and why databases still matter

    1:11:59
    Andy Pavlo1:12:00

    You can imagine a crazy world where you say, "I don't need a general-purpose data system anymore. For every single application, I want to vibe-code exactly a data system that does this for this one thing," and it can then be hyper-specialized. You kind of do this now with code generation or just-in-time compilation for some aspects of queries. And some systems do that: ClickHouse, Postgres, Umbra from the Germans. But hyper-specialization and making that be sustainable might be a bigger research question going forward.

    Matt Turck1:12:22

    So it's this, I don't know if it's a paradoxical kind of situation, but on the one hand, it's a bit of a stagnant industry right now in terms of evolution. At the same time, as we've hopefully established through the conversation, the layer itself is as important as ever, which is one of the reasons why your friend Larry Ellison is the—

    Andy Pavlo1:12:23

    He owns a Hawaiian island, right?

    Matt Turck1:12:25

    Always close to the richest man in the world, or was at some point.

    Andy Pavlo1:12:28

    He was. As of today, he's back down to eighth.

    Matt Turck1:12:34

    Oracle is a lot more than just a database company, but it's still the core product.

    Andy Pavlo1:12:36

    Headquarters is in the shape of a database.

    Matt Turck1:12:44

    Yes. But if nothing else, that shows the fundamental importance of databases in the—

    Andy Pavlo1:13:10

    But I would say no one is trying to make a serious attempt to replace Linux, right? Yes, there's niche operating systems for different environments, different hardware, and things like that, but no one's trying to build a significant-scale replacement for Linux entirely. FreeBSD predates Linux, from the BSD era from Berkeley, so that's why it's still around. But for databases, there's always new ideas where people think they can do things better than others. So I wouldn't say it's stagnant.

    Andy Pavlo1:13:36

    I would just say that a lot of the energy is going to be focused on Postgres in the short term, and there's certain things we can do to improve it. But meanwhile, I still think there'll be these sort of specialized engines. And I don't know what the next workload is. Again, it's not clear to me agentic workloads are significantly different, right? Because they're modeling things that are how humans design applications. Like the vector lookups, that's significantly different than what was in the past.

    Andy Pavlo1:13:50

    You needed additional operators. You needed additional indexes to do those things. It's not clear whether—I don't know what the next thing is, but there will be something. There always is.

    Matt Turck1:14:08

    So maybe to finish on that: the next five years, we don't know necessarily what that looks like. But in the meantime, does that mean the current companies, the ClickHouses of the world, just keep getting better and established because the fundamental need exists, and there's not a new wave of smaller startups?

    Andy Pavlo1:14:28

    So when you say better, though, there's two notions of better. There's, like, can an existing query run faster? Yes, ClickHouse has room for improvement. We can improve that. Other stuff is going to do the same thing. But like I said, the architecture, the high level of what ClickHouse is doing, is similar to what Snowflake did in 2013. Databricks is basically copying some of the existing ideas as well. Everyone is kind of doing the same thing, at least for analytics.

    Andy Pavlo1:14:54

    And what really matters, and one of the things that Snowflake did right, was part of the success is the stuff around the database system, like the user interface, developer experience, building to ingest data, interoperate with other things. That part they did really well. So it's almost like the things around the scaffolding of the foundation of a database management system that actually matters a lot. And one of the things that ClickHouse is actually—this is not research for me—but this is one of the things they are doing better.

    Andy Pavlo1:15:26

    Over time, it helps expand the reach of what you can use the system for. So the core kernels of the system can certainly be improved, but the fundamentals of what ClickHouse is doing versus other systems, at the smallest level, they're pretty similar. But the stuff around it matters a lot. And I think ClickHouse does some things well, can do things better. Snowflake does some things well, can do things better. Same with Databricks and the other guys.

    Matt Turck1:15:37

    Well, thanks for spending time with us. Congrats on the newish job that you also described in a rap analogy, right? What was it?

    Andy Pavlo1:15:54

    It's like when you look at Run the Jewels, it's like taking El-P and Killer Mike. Those guys were established on their own. You put them together, it's like people putting peanut butter and jelly together for the first time, right? So I like to think me and ClickHouse getting together is like that.

    Matt Turck1:16:01

    Okay, well, this is now officially on the record the most hip-hop-reference-heavy episode of The MAD Podcast.

    Andy Pavlo1:16:22

    I'm glad you found that in there in the blog article. That was my—it wasn't a test. That was my—because I wrote the blog article, I was floating it out there to see what ClickHouse would be okay with. Obviously I didn't want Larry Ellison in there, but I wanted to see, like, hey, if you get Andy, you're also getting other baggage.

    Matt Turck1:16:24

    All right, Andy, this was fantastic. Thank you so much.

    Andy Pavlo1:16:24

    Thanks, Matt. Happy to help.

    Matt Turck1:16:46

    Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.

    Andy Pavlo1:16:46

    Bye.