Yes, but Postgres is the best of breed now in many cases. So betting on your front end as a Postgres deployment that is compatible with Postgres, that might overcome that issue.
So, to talk about other parts of the database market, we talked about vector databases. Is that thumbs up or thumbs down?
What is thumbs up, thumbs down?
Like, in terms of, is that a good business to get invested in, or are they going to be around now that everybody else and their brother has a vector search capability?
One of the things that would happen before when I was a professor is, you always hear rumors about who's doing well, not doing well, through a combination of either the investors or former employees or students that maybe go to internships or whatever, or maybe interview some places and they come back. So you get sort of bits of information from everyone. You kind of piece together what the data landscape looks like. Problem is, now we're in ClickHouse, now I see everything.
As an investor, you see everything too. The one vector database company that I know is doing very well is Turbopuffer, and they are hyper-specialized in doing vector search at a cost-performance ratio that's much better than everyone else. So I don't think that the vector databases are going to go away. I think that they'll evolve in two ways. They'll have to become either a sort of general-purpose system like a Postgres, like a MySQL, where they become the system of record where you're storing the original tuples plus the embeddings or the vectors for them.
Or they become like an Elasticsearch, where there's a separate system where they have a copy of the data that's being pulled from the operational side. And in that case, they can live sort of comfortably as being this additional thing you add on. And if you want the raw best performance of vector search, in some cases you may have to go to one of these specialized systems. So I don't think that's going to go away. I just don't think I've seen predictions that, like, oh, Postgres is going to die at the hands of a vector database.
That's not happening. That's not happening.
Very much the opposite, right?
Graph databases we mentioned at the beginning. So, not to pick on them, but Neo4j has been around for 20 years now.
Yes. And this was supposed to be the moment for graph databases: AI. So what's happening there?
I have an outstanding bet with somebody on Hacker News where they said that by the year 2030, the graph database market was going to be larger than the relational database market. And if this becomes true, then I will wear a shirt that says, "I love graph databases," and I will use that as my driver's license, my university ID. I'll put it on my website till the day I die. I'm pretty comfortable. It's 2026.
We got four years to go. This is not happening. No, it's always been a niche market. And I think that, because my perspective on the research side, the research shows that if you do certain things in implementing the engine, which ClickHouse does do, DuckDB does some of this as well, there's things you can do that allow you to do the traversals of graphs, which are essentially just joins, self-joins on a table. You can implement those things very, very efficiently, and you can easily outperform Neo4j.
And that's kind of a kicking—like Neo4j, saying you're faster than Neo4j is like saying I'm faster than somebody that's in a wheelchair, right? You can run fast. It's a low blow. So, but I'm just saying that all the graph databases, I think even the best ones, you're just not gonna—
There's no system that has as much of these optimizations that are in the research and actually appearing in some of these systems now. You're gonna lose. What you will lose against a graph database is if you're doing the graph traversal with the client side and the server side, meaning, like, I got to figure out what the next node I want to go look at. I go back to the client and that decides the next node to go traverse.
If you're doing that back and forth, yeah, they'll beat you guys. But like I said, the SQL standard now supports property graph queries. Oracle has this, right? They were a big pusher of this. This extension of SQL allows you to do that traversal on the server side. So graph databases, I think, are a horrible idea. There's no reason why anyone would actually want to use them.
I know you have a special interest there. There was a cycle when the generation appeared, then went away. There seems to be a renewal. What is a GPU database, and what is your prediction?
So, a GPU database is a data management system where the execution engine for queries is offloaded to a GPU running on PCIe, or running in the same box or another box. So, the history of people trying to build accelerators for data systems goes back to the beginning of data systems. In the 1970s, they were called database machines. So people would build specialized hardware to run sorting and query execution operators. And that obviously died out in the early 1980s because by the time it'd take you to design and fab new specialized hardware, Intel or Motorola would put out the next CPU, or the hardware got better and just the gains you were getting went away.
So hardware accelerators for databases basically died out in the 1980s. There wasn't a lot of activity in the '90s, 2000s. You saw sort of the rise of people trying to do FPGAs for databases, and every so often that comes back now. Some of the cloud vendors do a little bit of these things, but usually to filter things on the NIC, on the network side of things coming in. So, with GPU databases, again, there was a bunch of systems in the 2010s that were trying this.
We did a seminar series at the university where we invited all the GPU database guys to come and give talks about what they were doing, why they were faster than existing systems. And the big challenge at the time was that those systems, you had to put the entire database inside the memory of the GPU. Because if you had to go back up through PCIe, it was just way too slow. And a bunch of those startups sort of fizzled out. Some of them are still around, but they're sort of specialized for doing visualizations.
And then there was, in the last year or so, NVIDIA has basically gobbled up a bunch of these GPU database companies that were kind of struggling along. And I was an advisor for one of them called Voltron, but they also picked up HeavyDB. And so NVIDIA is all in on this now. So it remains to be seen whether the idea that you're going to build a CPU-only database system, long-term, whether that's going to still hold. I've heard, again, mixed reports. This is public.
Microsoft has offerings now in the cloud that can be accelerated for your data system, can be accelerated with GPUs for analytics. Another major database company that I can't say who they are, they looked at the economics of GPUs and decided it wasn't worth it. So one of my former students, now a professor at the University of Wisconsin, they're now on leave at NVIDIA. They have a project called Cider, which is not necessarily a new database system, but it's a layer in between an existing system and CUDA.
And so it supports taking DuckDB queries and running that down on the GPU. I think they can do this in Doris or StarRocks and DataFusion. And so at ClickHouse, we've been potentially looking at this as well, but it's research. I don't know. It's interesting to see how much translation you have to do between how ClickHouse expects things and how CUDA wants things to be, the data layout and so forth. How do you organize memory or share memory between these different components?
TBD. It remains to be seen whether this actually makes sense. But certainly, there's a lot of research energy behind this. And publicly, I can say this: NVIDIA is obviously pushing this because it'll sell more GPUs, right? Because it's hard enough to get new CPUs. Everyone's compute-bound, or memory is hard to get. The computing hardware is very expensive, hard to get now. And GPUs, of all the things, are the most expensive hardware to get. And now you're going to say your entire database is going to run off GPU.
I don't know if that makes sense, at least in the short term. But if the performance improvements are quite significant, and some of the research shows that it is, maybe it makes sense.
Is there an emerging category, or maybe a niche somewhere within a category, that people don't talk about enough yet?
I mean, you can always build new data systems for new hardware, but my track record on this is terrible. We've done a bunch of research on experimental hardware, and it always gets canceled. Or it's even not that experimental. It's like Intel had this Optane persistent memory stuff. We did a bunch of research on building systems for that. Because if you assume now your DRAM is persistent, like you pull the plug and you don't lose anything, that changes how you fundamentally build a data system.
We did a bunch of work on that, and then Intel killed that product line. We were doing other research on processing-in-memory hardware. So think of DRAM sticks with CPU cores directly on the DIMM. So the data system now can say, okay, instead of pulling things from memory, bringing it into my CPU caches, and then I can compute things on them, I'll just send the query down to the DIMM itself and run it there. We were doing a bunch of work on this thing called UPMEM that got bought by Qualcomm and got killed last year.
So that didn't work out. So there's always a bunch of work you can do on data systems, on new hardware. I would say that, actually, I don't know the answer, right? This is one of the things I'm trying to figure out at ClickHouse. And we talked about this in the very beginning: are agentic workloads significantly different than what humans or what existing applications do now? And if so, why or how? And how would you change maybe the development of a database system to take better advantage of this?
That remains to be seen, how that works. I think there's always a bunch of problems in query optimization that I think are interesting. That remains the hardest part about database systems. Incremental materialized views, another big challenge. Again, these are not things that no one else has thought of. People have been trying to do these things for decades. So I think the agentic stuff is probably the most interesting and relevant thing to me right now: what changes with these workloads, and what changes in the system architecture?