The Death of Big Data and Why It’s Time To Think Small | Jordan Tigani, CEO, MotherDuck

The MAD Podcast with Matt Turck · with Jordan Tigani, CEO, MotherDuck

Jordan Tigani is the CEO at MotherDuck. We cover why most big data workloads query only a hot slice of archived data, how distributed systems can add roughly 400 milliseconds of overhead where local queries run in single-digit milliseconds, and why MotherDuck splits query plans between local DuckDB instances and the cloud for 60-frames-per-second interfaces.

Watch on YouTube

Chapters

  1. 0:56 — What is the Small Data?
  2. 6:56 — Marketing strategy of MotherDuck
  3. 8:39 — Processing Small Data with Big Data stack
  4. 15:30 — DuckDB
  5. 17:21 — Creation of DuckDB
  6. 18:48 — Founding story of MotherDuck
  7. 24:08 — MotherDuck's community
  8. 25:25 — MotherDuck of today ($100M raised)
  9. 33:15 — Why MotherDuck and DuckDB are so fast?
  10. 39:08 — The limitations and the future of MotherDuck's platform
  11. 39:49 — Small Models
  12. 42:37 — Small Data and the Modern Data Stack
  13. 46:47 — Making things simpler with a shift from Big Data to Small Data
  14. 50:04 — Jordan Tigani's entrepreneurial journey
  15. 58:31 — Outro

Transcript

What is the Small Data?

Matt Turck [1:28] It feels like the absolutely inescapable, unavoidable way to start this conversation is to talk about small data. Just when we thought we had finally made it in the world of big data, you came up with a very well-written and very noticed blog post in early 2023 called "Big Data Is Dead." And you built a whole thing around this. And just last week in San Francisco, you ran the Small Data Conference. So, small data, what is it all about?

Jordan Tigani [1:54] So I'm very glad you said that it was unavoidable. We did a lot to try to get the message out and get people excited. And it feels like it's a little bit sort of counter to the prevailing narrative that everything for 15 years has been about big data this, big data that. How big is your data? How much can you scale? And kind of the tipping point for me was when I saw the Databricks-versus-Snowflake benchmarking war, and everybody was focused on the war between Databricks and Snowflake.

Jordan Tigani [2:39] And to me, the biggest thing that I noticed was, while they're looking at the database benchmark, the query sizes they were using were 100 terabytes. And I remembered back from my time at BigQuery, we had some of the largest customers in the world. We had Walmart, Home Depot, Equifax, HSBC, and nobody was running queries anywhere near that because it would have actually fallen over at the time. And so I knew that people were kind of not even really pushing up against limits.

Jordan Tigani [3:14] And so I thought, wow, if this is sort of the state of the art that people are focusing on, this size of data that nobody has, there's got to be an opportunity actually to look at the smaller data sizes. And I kind of was remembering back to when I was doing a bunch of analysis also when I was at BigQuery on the query sizes that people were using. And most people actually had small data. The amount of data they were actually using was even smaller than that.

Jordan Tigani [3:34] And so I was kind of like, hey, I bet if you were going to design a system these days from scratch, you'd do it differently. After Google came out with MapReduce and GFS and Bigtable, kind of everybody's—

Matt Turck [3:36] All of which was in like 2006, right?

Jordan Tigani [4:07] Yeah, 2004, 2006. Everybody's brain just sort of broke, and they're like, wow, in order to build systems that can handle the data sizes that we're seeing, you have to just dramatically change how you're building them. You have to run on lots of cheap, inexpensive machines versus these giant, hyper-expensive machines. And to be fair, that was a problem at the time. But nowadays, I've got a Mac M2 laptop. It's two years old. It's probably an order of magnitude to two orders of magnitude faster than the server machines were back when MapReduce came out and people started building these systems, let alone nowadays the server systems have hundreds of cores and can have terabytes of RAM.

Jordan Tigani [4:46] And so really, if you were going to build something now, why would you bother with all the complexity of scale-out? Because the thing about the way we design systems—I was one of the people that helped start Google BigQuery, and I worked on SingleStore for a couple of years, so I have been with my elbows deep in building these complicated systems—is that there's just this huge tax that you pay to build a distributed system that scales out and can do distributed transactions and shuffle data.

Jordan Tigani [5:20] And if you were going to design something for modern hardware, you could make it much, much faster, and you could make it much simpler, meaning you could actually progress faster. Because that was one of the things: there were a couple of well-known join optimizations that we added to BigQuery, and they took like a year to do. And they weren't that hard; it's just that getting everything right meant that it took a long time.

Jordan Tigani [6:07] So, sort of getting back to small data, the idea is that if you have massive amounts of data, you have to build these complex systems. If you have smaller amounts of data, you actually don't need such complexity, and you can move faster, less expensively. And also, now that networking speeds and local machines, laptops, have gotten so much more performant, you can actually push workloads down to the end user. And that opens up so many new and different architectures and different ways of handling data and building systems.

Jordan Tigani [6:41] And I think we're starting to see that with some startups out there that we invited to this Small Data Conference. The conference was sort of born—it was the idea of Bob van Luijt, the CEO of Weaviate, and he's like, hey, what do you think about doing a small data conference? And I'm like, that would be—I just love the idea, because we try not to take ourselves too seriously at MotherDuck. And you have Big Data London and Big Data this and Big Data that, and we could just sort of poke a little bit of fun at those things and do things a little bit differently.

Marketing strategy of MotherDuck

Matt Turck [7:29] And I was, like, looking at the manifesto that you have on the Small Data SF 2024 website. I love it. Clearly, you're having a lot of fun. Just from a marketing perspective, I have to give you kudos. It's amazingly well done. It's so hard, if you're a technical company, to break through the noise and come up with something that feels like a movement and feels like a manifesto. And you guys have done a remarkable job doing that. And yeah, it's really fun to watch.

Jordan Tigani [7:49] Thanks so much. I think a couple of times in my career, I have built things that I thought were amazing technology and didn't get the other pieces right, and nobody saw them. And so distribution matters, as it turns out. Yeah, distribution excitement. And so, when we started MotherDuck, it was very deliberately to make sure that we weren't just writing a bunch of code and throwing it over the wall and saying, "Hey, if we build it, they will come."

Jordan Tigani [8:29] That we were also trying to tie it to some thought leadership, take advantage of—we have several people in the company who have a lot of experience building these kinds of systems and had seen the trend going in this direction and where we thought the world should go was in this other direction. And so, had something we could be passionate about and that we could write about and add a little bit of fun and sense of humor to it, I guess, always helps.

Processing Small Data with Big Data stack

Matt Turck [9:07] So, a couple of just questions which I'm sure you've gotten a thousand times. If you already have BigQuery in place, or Snowflake, or Databricks, or the whole big data kind of infrastructure, or modern data stack, whatever you call it, if you can do the big data stuff, can't you do the old data with it? I mean, shouldn't they be priced in a way that ultimately that should not make a difference to you?

Jordan Tigani [9:16] Yeah. So I think I mentioned before that there's this tax you pay with these complex distributed systems, and you pay the tax twice. You pay the tax in terms of latency.

Matt Turck [9:23] It's just, when you think about it, latency is going to query the whole thing.

Jordan Tigani [9:52] Well, you're going to query the whole thing, but also, when you send a query to BigQuery, for example, your query's gonna get spread out over possibly a thousand or thousands of machines. All of that coordination takes time. There's this scheduler allocating the slots. There's all the RPCs dealing with canceling things that had been running on them previously, getting all the results, aggregating the results. There's just a limit to how fast you can do that versus if it's just the first machine that you hit actually runs the query. You can do things—I mean, we can do queries in sort of single-digit milliseconds.

Jordan Tigani [10:25] And in BigQuery, we were very, very happy when we got the overhead down to like 400 milliseconds. So that's like two orders of magnitude difference. And then on the other side, it's just the cost tax. To have all this hardware creates a lot of overhead. And I think a big bank did some benchmarking against BigQuery, Snowflake, and SingleStore when I was at SingleStore, and they estimated that on a per-core basis, BigQuery was 40 times less efficient than SingleStore, which is kind of a more dedicated, performance-optimized query engine.

Jordan Tigani [11:15] But so you're giving up at least an order of magnitude. And for Google, it's like, whatever, we own the hardware, we'll just throw lots of cores at it. But at some point, you got to pay for those. Somebody's got to pay for those cores, and somebody's got to pay for that inefficiency. And so I think if you build these systems more simply, then you can have dramatically less expensive. And then there's another bit that people don't always recognize, which is sort of the velocity of improvement.

Jordan Tigani [11:47] And often, when you choose a technology, you want to stick with that technology for at least a few years. And so really what you're buying is not just where it is now, but where it's going to be a year from now, two years from now, five years from now. And when you have these complex distributed systems, they get better very slowly. And if you look at DuckDB, for example, which is what my company is built on top of, incredibly rapid pace of improvement, incorporating brand-new stuff coming from algorithms, join order optimizations coming from academia. The graduate students that come up with those implement them in DuckDB, and then they ship them, and then they're out like a month later.

Jordan Tigani [12:27] And so I think the rate of improvement is also something that you have to take into account, that some of these big data systems are just going to have a hard time keeping up.

Matt Turck [12:57] Is some of it a question of use case as well? Meaning, if you use data, as in modern data stack, for purposes of BI, effectively data analysis, then maybe small data makes a lot of sense. But if you want to, I don't know, train a big machine learning model, then, as abundantly documented, you want as much data as possible. Is there some truth to that?

Jordan Tigani [13:17] Yeah, absolutely. There are clearly some use cases and some workloads that are not small data. I think big data may be dead, but it's not going away. And I think the one thing that I often hear, though, is sometimes people are like, "Well, I read Big Data Is Dead and I agreed with most of it, but I've got big data." But I think actually, because there's a lot of people that have a lot of data, what they actually use is a small section of data.

Jordan Tigani [13:50] So you might have 10 years' worth of logs. It might be a petabyte worth of logs, but if you only actually query the last seven days or the last day, that's not really big data. Because of separation of storage and compute, that other data just sort of sits there and is cold on AWS S3, or on object storage somewhere. The important part is the part that you're querying. And the vast, vast majority of workloads actually just query that hot data because it's expensive to query the whole thing.

Jordan Tigani [14:24] I mean, I used to run this query. I used to give talks on BigQuery, and I'd get up on stage and I'd say, "Hey, look at this dataset, it's a petabyte, and I'm gonna query this whole petabyte." And I hit Enter, and it would query the petabyte, and I would be like, "Look, isn't that amazing? I queried a petabyte." And the thing that I didn't tell you is that it cost $5,000 to query that petabyte. And I think big data has been able to make it possible to do some of these giant things.

Jordan Tigani [14:58] But if you still have to do all the work, you still have to do all the work. And they haven't been able to make that inexpensive. And so I think part of the small data is actually recognizing that, okay, in order to not have massively expensive analytics bills, you want to actually slim down, aggregate the work that you're doing. And often, people have the medallion architecture. You land data in the bronze tier, and then you transform it a little bit into the silver tier.

DuckDB

Jordan Tigani [15:30] And then finally, you have your gold presentation tier. Very often, that presentation tier, that gold tier, is pretty small. And I think, as you were mentioning, the stuff that you run your BI off of, if you have a human waiting there, you want that to be fast. And that's just kind of another reason to use a low-latency, small-data tool to operate on the gold-tier data.

Matt Turck [15:37] So I'd love to go into MotherDuck specifically. So what is DuckDB?

Jordan Tigani [16:06] DuckDB, like SQLite, is just a library. It's something that, as you're building your code, you link in. It's functions that you can call inside that library, and you don't have to set up a separate server somewhere and call out to that server. It just sort of runs everything inside your process. And from a latency perspective, that's super nice. From a complexity perspective, it's also super nice. If you're running in Python, for example, the way Python works is kind of, things in Python have access to the variable namespace.

Jordan Tigani [16:44] And so actually, in DuckDB, you can query against your Python objects. So if you just import DuckDB in your Python process, you can just start querying data without even having to move it or prepare it or do anything. So it's just really, really high developer experience for accessing your data. It also works well as a sort of standalone database as well. It's not just in-memory, and because it's just a library, it has no dependencies, and it's incredibly lightweight.

Jordan Tigani [17:19] It can run in the browser, so it can run under WASM in the browser. So you can have a full-blown query engine running in your browser. It's like, wow, I can run a bunch of SQL queries right in my browser without installing anything, which is sort of, I think, one of the things that makes it pretty unique.

Creation of DuckDB

Matt Turck [17:27] And historically, that was a project developed at a university?

Jordan Tigani [17:34] At a research institution in Amsterdam called CWI. It's also where Python was created.

Matt Turck [17:37] And so who created it? Who are the creators of it?

Jordan Tigani [18:05] So yeah, it was Hannes Mühleisen and Mark Raasveldt. Hannes was a professor, Mark was one of his graduate students, and Hannes had just gotten tenure, and so nobody could tell him what to do for a little while. And Mark had finished his PhD papers early, and so nobody could really tell him what to do for a while. And they're like, hey, we've been using MonetDB, and there's a bunch of limitations. Let's write our own to solve some problems that they'd seen in the data science world, and they're like, data scientists hate databases.

Jordan Tigani [18:36] And partly it's because they had to install databases and configure them and load data into them. And they said, hey, there's a better way. And it turns out that they're amazing database researchers, but also great engineers, and they were able to build a super useful system that just started becoming more and more popular.

Matt Turck [18:39] And when was this? What, roughly what year?

Founding story of MotherDuck

Jordan Tigani [18:48] I think it was 2018. I think they've been going for five or six years. I think it's been like five and a half years now, so probably 2018.

Matt Turck [18:52] Tell us about where MotherDuck fits in the puzzle.

Jordan Tigani [19:19] I saw a tweet while I was still at SingleStore, and somebody was doing some benchmarking of SingleStore against BigQuery and Redshift and this thing that I'd never heard of called DuckDB. And I'm like, wow, that's really fast. Where did this come from? How is that possibly that fast? And I started digging into it, and I'm like, oh, it's a research prototype. And then I realized, I actually kind of looked at it, and I saw their blog post on time zones, and in BigQuery, we didn't implement time zones for, I think, like six or seven years because it's just really hard to get right.

Jordan Tigani [19:47] And there's all kinds of bizarro things when you deal with time zones. And we're like, well, we don't have to worry about that. Everybody can just convert from UTC, from Greenwich Mean Time. And the fact that they actually had the attention to detail to do this, it was like, hey, there's something going on here. This is not just something that you write, that you build to sort of do a couple of papers on to get your PhD.

Jordan Tigani [20:34] This is real. It scales down so small, it's so lightweight, you could build an amazing serverless system using this, and somebody should put it in the cloud and build a service around it. And I'm like, hey, I helped start BigQuery. I've done one of these before, and I helped with the SingleStore SaaS service. And I've done a couple of these. I kind of know how it works. Maybe that should be me.

Jordan Tigani [20:48] And I had never really thought of starting a company before, but the idea just was so compelling, and the technology was amazing. And I met Hannes and Mark. I was actually on a—

Matt Turck [20:51] So you just emailed them and said, hey?

Jordan Tigani [20:58] Yeah, I got an intro from a mutual acquaintance, actually Lloyd Tabb, the founder of Looker.

Matt Turck [21:02] In the world of common acquaintances, that would be a good one.

Jordan Tigani [21:24] I knew that he's working on Malloy, this metrics layer thing, and I knew that that would work with DuckDB. And so I asked for an intro to those guys. He's also like, yeah, while you're at it, you should talk to my friend Tom. He invested in Looker. He'll give you some feedback on the idea.

Matt Turck [21:36] Tom was Tom Tunguz, who we did a great episode of The MAD Podcast with, which we'll, I guess, add to the show notes, as real YouTubers say.

Jordan Tigani [21:50] Yeah. And he did our seed for MotherDuck. And at first I was asking for a job. I'm like, hey, what you guys are doing is really cool. I'd love to come. If you guys are going to build a cloud service, I'd love to come work with you on this.

Matt Turck [21:55] Clearly you're going to do that, right? Why wouldn't you do that? Yeah, sure. And they said no. They said, no, we don't want to do that.

Jordan Tigani [22:21] They said, no, we just really want to focus on kind of building the core database, and we're not really interested in that. But if you were going to do something like that, we'd love to partner with you. And that sort of got me thinking. And I was—my first vacation post-COVID, I was in Portugal, and then I kind of rerouted the return. I went through Amsterdam instead of through Paris to meet Hannes and Mark. And they'd scheduled like four hours in the afternoon.

Jordan Tigani [22:48] And I'm like, we're a bunch of nerds. What are we gonna talk about for four hours? We're gonna talk for like 45 minutes and start getting nervous and awkward. But we just started talking, and we were just sort of geeking out about this database and that database, and next thing we know, it's dinnertime. And we went to dinner with my wife was there and Hannes's wife was there.

Jordan Tigani [23:07] And so it was just this really cool meeting of the minds, and we realized that we could work together, I think. And yeah, so that's how we kind of started working together. And I think what we did is, I'm hoping that if this works out, and I'm obviously hoping that it works out, that it will be kind of a model for how to do an open-source-plus-corporate-VC-backed arrangement, because they have a foundation that kind of owns the DuckDB IP.

Jordan Tigani [23:47] It's MIT-licensed. It's always going to stay that. It's always going to stay kind of open source. But we also gave them a chunk of the company when we started, essentially a cofounder share. So they're incentivized for us to be successful, but also we have no direct say on anything that they do. And Hannes may decide to go build something weird or something that we don't like.

MotherDuck's community

Jordan Tigani [24:10] And that's just how it works. On the other hand, he did agree not to work with anybody who's doing something similar to what we're doing. But he's also incentivized not to do something kind of too—

Matt Turck [24:30] What about the community? I mean, I guess the community does what the community does, but it's sort of, what's the word, like a sort of common wisdom in startup and venture circles that the commercial company building on top of the open source should also kind of own the community, to the extent that any community can be owned. How does that work for you guys right now?

Jordan Tigani [25:05] We kind of have—they're somewhat disjoint communities. We have our MotherDuck community, and we have, like, a Slack. But also, there's a pretty vibrant DuckDB community that we have not jumped in and tried to own or run anything. I don't think that would've gone over well. On the other hand, we try to be helpful where we can. We contribute a lot back to DuckDB code, and we have great relationships with the founders and with the community. And so I think that it's actually kind of a happy way of doing things.

Jordan Tigani [25:24] I mean, I think DuckDB is super popular, and we don't want to try to kind of horn in on that as long as we can be successful by building our managed service.

MotherDuck of today ($100M raised)

Matt Turck [25:36] And tell us about the company today. So I think you raised about $100 million across three rounds. Is that the right number?

Jordan Tigani [26:02] Yeah, we raised our seed round from Redpoint and Madrona and Amplify, and then we got preempted a few months later by Andreessen. And then we got preempted for our B about six months later by Felicis. And building a database as a service is expensive. DuckDB is an amazing piece of software, but it's not a data warehouse. And to turn that into a data warehouse is hard. There's a lot of things that, if you poke it the wrong direction, it'll fall over.

Jordan Tigani [26:36] And then it's sort of like, well, nobody's ever poked it in that direction before. And so it's just stuff to work through, stuff we have to build. And we're building this, I think, pretty rich database as a service and the serverless backend, this highly multi-tenant, this hybrid execution system, or actually dual execution system, where we can push workloads down to the end user and split query plans. And so I think we're doing some interesting, non-trivial stuff, partly because there's a model that I think I've seen before, which is in open source, where somebody builds this great open-source project, and then that's what they completely focus on because, to get adoption, to get excitement, to get funding, they have to build a great open-source project.

Jordan Tigani [27:36] And then they say, well, we'll monetize with SaaS and we'll put it in the cloud. And it's sort of an afterthought. They have a couple of junior engineers work on it. And so you're basically just running the thing in a Kubernetes container, and it ends up being something that's kind of trivially clonable by AWS or Google or somebody else. And you'd end up without actually much of a moat. And to me, actually, that is disappointing, not because the hard work gets cloned elsewhere.

Jordan Tigani [28:10] But it's disappointing because I think that there's a lot of things you can do, a lot of interesting architectures that you can build if you really focus on the SaaS service. And so for us, for MotherDuck, we are focusing on building this kind of, I think, unique way of delivering the service to users. And then DuckDB Labs gets to focus on building this great, amazing database. And because our API is not a typical sort of web API, you send a query and get a result back.

Jordan Tigani [28:45] Our API is like this partial query plan API where you send a piece of a query plan and you get a piece of a query plan back. And it's complex, but it's complex because it can actually deliver value. It can deliver value in that you can join data from your Postgres server against data that's living in the cloud, build UIs that let you query data at 60 frames per second, and fly through your data as if it's a video game. And the reason that we can do that is because we can do part of the work locally and we can do part of the work in the cloud.

Jordan Tigani [29:05] So, a long-winded answer, but we are trying to spend a lot of our innovation tokens on how the SaaS infrastructure delivers a service to users.

Matt Turck [29:14] You mentioned, as we were saying right before we started recording, that you have four offices: Seattle, New York, Amsterdam, and SF.

Jordan Tigani [29:45] Yes. So, yeah, we have about 50 employees. We kind of started during COVID, and everybody was explicitly remote. But I think, like a lot of people, we recognize that there's value to actually having people in person and collaborating and just being able to feel like you're part of a real team. And so we realized almost everybody was in one of these four cities. Seattle—I'm in Seattle—that's where our headquarters is. San Francisco is just a huge concentration of talent.

Jordan Tigani [30:10] New York, a couple of co-founders were there. Plus, it's sort of a nice balance between the West Coast and Europe. And then finally Amsterdam, because we wanted to be close to the DuckDB team, so that we can kind of go over and poke and say, "How's that coming?" Or take them to dinner, or just—

Matt Turck [30:12] Can you do this faster?

Jordan Tigani [30:52] Help us actively manage the relationship. Actually, we're pretty evenly distributed so far. And I think it's working pretty well. Each office is kind of developing its own flavor. And our New York office is probably the youngest office and the most fun. Of course, it's New York. Amsterdam is the most academic. But yeah, we raised $100 million, and we have several thousand users. It's growing super fast. We're deliberately pricing it pretty low. We want you to be able to get in and get a data warehouse that works for $25 a month.

Jordan Tigani [31:23] And enterprise prices will, of course, end up being more than that. But we think you shouldn't have to spend thousands of dollars a month to do basic data warehousing and analytics.

Matt Turck [31:27] And do you want to explain then what in-memory means?

Jordan Tigani [31:57] So, transactional databases, typically you are operating on sort of one thing at a time. You have an order and you create an order, and maybe you update the state of that order, and there's a bunch of consistency checks to make sure that that order matches a real customer and matches a real product and line items, et cetera. And analytical databases, on the other hand, tend to operate across data. So you can ask questions like, well, how many orders did I have in the last week?

Jordan Tigani [32:33] Or how many orders broke? Who is my biggest customer by amount that they spend, and then broken out by region? And those types of questions—the data tends to be stored differently in these types of databases: column store versus row store. And then there's also a subclass of databases, which is sort of an in-memory database, which means that if you turn your computer off, you lose all the data. But on the other hand, memory is four orders of magnitude faster than disks.

Jordan Tigani [33:04] Depending on what kind of disks, but it's generally much, much faster. And so you can do things blindingly fast. And I think there was a wave of in-memory databases that started about 10, 15 years ago. MemSQL, which became SingleStore, which is where I worked previously. But I think it turns out that people really want persistence. If they take all the energy to load this data into their database, they want it to be there and they want it to be transactionally updated, et cetera.

Why MotherDuck and DuckDB are so fast?

Matt Turck [33:23] What makes DuckDB and MotherDuck then so fast and so appropriate for those use cases?

Jordan Tigani [33:47] One of the things that makes it fast is it's a brand new database built from scratch, applying the latest and greatest best practices. There was a paper that Michael Stonebraker wrote. Michael Stonebraker is a Turing Award winner, like the Nobel Prize of computer science in databases.

Matt Turck [33:49] And creator of many database companies.

Jordan Tigani [34:22] Yeah, he created Postgres and Ingres. And I think the paper was called No Free Lunch. And it was really about that, like, hey, the technology has changed dramatically. Why haven't databases changed? If you just think about SSDs, like you might have in your laptop, they don't have the spinning platter anymore. And because of that, they're much, much faster to find things. There's some things that they do that are dramatically better. There's some things they do that aren't as good.

Jordan Tigani [34:53] But if you were going to build a new system, you'd just build it differently. But people ended up like, well, the old one works kind of well. And it's a decade or even more since that paper. More has changed in hardware, in sizes and speeds. You get these, what are they called, almost memory-speed disks. They totally upend the way people tend to think about the trade-offs in software.

Jordan Tigani [35:24] It used to be if you had to hit disk, it was like all of a sudden your program just screamed to a halt. And nowadays it's not so bad. You can amortize the costs of disk access. So DuckDB is built with modern techniques, built really well by some really brilliant people, and made very fast. And I think one of the other things they did is they avoided traps that kind of prematurely optimized.

Jordan Tigani [36:01] So one of the things people talk about for analytical databases is vectorized execution, which is not to be confused with vector databases. Those are actually something different. But in vectorized execution, you basically take a whole bunch of inputs, a vector of inputs, and you can process them all at once. And this is very nice for how computers use caches. It turns out to be very efficient to write the code for those. But typically what people do is get so excited about it that they're like, okay, I'm going to write special machine instructions that do this and do this fast.

Jordan Tigani [36:39] And we did this in BigQuery, and ClickHouse does it. A lot of these engines build these hand-coded vectorized execution engines because there are computer instructions that specifically deal with lots of things at once. But DuckDB wrote it in such a way that they let the compiler do that optimization for them, and they probably leave a little bit of performance on the table. But on the other hand, it just means that you don't have to maintain all this crufty, really complex, nasty, hairy assembly code.

Jordan Tigani [37:16] You just have these careful, elegant, elegantly written operators. And it also meant that, for example, when Apple Silicon came out, it took them like two hours to make it work on the new Apple Silicon, versus having to do all this complex hand-coding stuff. So there's a bunch of reasons why DuckDB is fast. MotherDuck is fast because we let DuckDB do its thing. We always have a DuckDB running on your client.

Jordan Tigani [37:47] So that could be in the web browser. There's a DuckDB running in your web browser as well as on the server. Or if you're running in a Python process or you're writing some code that connects to MotherDuck, there's a DuckDB inside that process. And that DuckDB, we can basically cache data there. And so queries that can hit that cache don't even have to hit the server at all. So if you think about, we have some UI apps that we've built that will start out and they'll have to run queries against the server, but then they'll end up building this cache. And then as you navigate the UI, everything happens locally.

Jordan Tigani [38:35] And so this is how you get the 60 frames per second visualization speeds that are literally impossible in more traditional architecture. Because if I'm talking to a data center that is on the other side of the country, that's 100 milliseconds. And so the best I could possibly do, even if it was infinitely fast to do the query, would be 10 frames per second. And so you end up being limited by the cloud architecture.

The limitations and the future of MotherDuck's platform

Jordan Tigani [39:09] But by pushing stuff down to the client, the only limits are the limits of the local execution speed. And the local execution speed is just much, much faster than it used to be, and getting faster. And the other thing is more and more people have fiber to the home or fiber to the workplace. And so pulling some of these datasets down, which used to be prohibitive, now you can do in seconds or less.

Matt Turck [39:17] Maybe to close on product, what are MotherDuck and DuckDB not good at yet that you have on the roadmap?

Small Models

Jordan Tigani [39:50] Scale is certainly one limitation. And I think if you're going to push past working sets of like 10 terabytes or larger, then MotherDuck doesn't work well yet. We are working on larger machine sizes and just sort of things that are going to scale better as you get bigger. But that's certainly something that could come up.

Matt Turck [40:12] You mentioned vector databases. So that's another inevitable question about where MotherDuck fits in the nascent AI infrastructure stack. I saw, again, on the website for Small Data SF that you were talking about small data, but also small models. Maybe walk us through that.

Jordan Tigani [40:40] As we're pushing, I think there's a real analog to pushing data workloads down to the client and then also being able to push models down to the client. And I think everybody who's worked on AI things has probably been sort of frustrated. You type something in ChatGPT and you kind of sit and you wait and you twiddle your thumbs and you wait. And yes, it's incredibly powerful, and you get these magical responses back, and they make it feel like it's faster by streaming results back, but it's actually quite slow.

Jordan Tigani [41:31] But if you can do a lot of the work, or certain types of questions, or certain pieces of questions via a local model, then you can get much more interactive AI applications. And especially if you're doing, like, RAG or something, where the retrieval-augmented generation, where you're pulling data from somewhere, you're using basically a database to store state, and then you're using an LLM to kind of understand something about the world. Now, if you do it that way, you can actually use smaller models because they don't have to quite understand, they don't have to encode all the state as well.

Jordan Tigani [42:21] And so you can do local RAG, where you're basically doing local lookups and local inference. And then perhaps for the harder things, you can call out to the server, or you can start with an answer and then refine the answer from something that is remote. But I see there are some really nice parallels between the architectural stuff that we're doing in MotherDuck and things like Ollama, where you can actually operate on models locally. And there's certainly really interesting things coming out with hardware and with compute environments that lets you get access to the GPU and do more AI stuff on your local client.

Small Data and the Modern Data Stack

Matt Turck [43:03] I saw that you seem to be good friends and partners with George at Fivetran, and I would have assumed that small data is not a good thing for the Fivetrans of the world—not to pick on them—but, like, any company in that stack, because don't you need a lot of data and a lot of complexity for this modern data stack, this suite of vendors, to thrive?

Jordan Tigani [43:35] I mean, it's interesting you mentioned George and Fivetran because he was a speaker at our Small Data SF conference. He and I did a town hall conversation, and one of the things that he said was, like, yeah, it's shocking how little data people use. And generally, if people are pushing a lot of data through Fivetran, it's because they're doing something really inefficient where they're basically just sort of—they're basically recopying all their data every day. But typically, the sizes of data that they see are much smaller than people would expect.

Jordan Tigani [44:08] But I do think that the modern data stack, the ideas behind the modern data stack, are important. And I think really that you have, you kind of have, like, I would call it sort of three, maybe four pieces. You have the data ingestion, you have kind of your query engine, and then you have your visualization layer. And then maybe the fourth would be kind of the orchestrator, and that would be sort of dbt. So, like, Fivetran would be ingestion, Snowflake would be the query engine, Looker the—

Matt Turck [44:14] Visualization layer.

Jordan Tigani [44:33] The visualization layer, and then dbt the orchestration. Obviously, you can swap out those pieces, and there's lots of competitors in those spaces. I think kind of the world is still arranged in those buckets. And so I think kind of that is still valid. Like MotherDuck, I think we're playing in the, hey, we have a great query engine space, and you can hook us up to your BI tool, and we connect to Tableau and Power BI and Omni and Preset, et cetera.

Jordan Tigani [45:14] And then you can hook us up to your ingestion engine, whether it's Fivetran or Airbyte, et cetera, and it works great in dbt. And so we're playing well in the ecosystem. Does that, fast-forward five years from now, does all that still look the same? Does the advent of things like these open data formats, does that start to kind of change how some of that work gets done?

Matt Turck [45:29] Does it impact you at all in a world where Tabular—and that was post the acquisition of Tabular by Databricks—and the rise, precisely, of this open data format? Does that change anything for you?

Jordan Tigani [45:54] It does sort of open up some doors for us. If you think about a lot of people who are—and we're hearing from a lot of people that are using Snowflake—and they're moving their data into, usually, Iceberg as part of a cost reduction, a way to avoid lock-in, a way to be more flexible and have more flexible access to their data. And that's great news for us because if the data is locked in Snowflake and we want somebody to try MotherDuck, well, we have to convince them to export the data or write it to two places, have multiple copies.

Jordan Tigani [46:32] And it's a mess. It's a migration. If the data is in Iceberg, we have the same access to it that Snowflake does. And so I think that's going to be hard for the incumbents, and I think it's going to be a net benefit to the people that are coming in with new tools and new ways of doing things. And it's going to put pressure on margins, which, again, tends to be in favor of people that are coming in afterwards, especially if they have a simpler architecture and can deliver things less expensively.

Making things simpler with a shift from Big Data to Small Data

Matt Turck [47:17] There is batch processing, which has its one stack. There is real-time processing, which has a different stack. And then there was Big Data, and now there's Small Data, and sort of, what do I do? It feels like things are getting more complex rather than less. But is part of the Small Data message that things are actually getting simpler, but you sort of need to get rid of some of the pieces of the past?

Jordan Tigani [47:45] I think one of the things with Small Data is, I think because the architecture is simpler, we can focus more on building better experiences. So yeah, maybe there might be a plurality of tools involved, or a plethora of tools involved. If those tools are simpler to use, then I think kind of the net cognitive load can go down. And I'll give an example: a lot of data is in CSV files. And as much as, sort of like, as a database person, it sort of makes your head explode.

Jordan Tigani [48:25] You're like, "Why would you put all this in a CSV file?" It's just, that's the way the world works. And it's simple, and it's easy to write a CSV parser, and it's easy to write CSV, but there's so much broken CSV out there because it's actually really hard to write a totally unambiguous, correct CSV file. And everybody sort of does it differently: different null characters, and is two empty quotes, is that a null or is that just an empty string? Like, there's just all sorts of weird things that happen.

Jordan Tigani [48:57] And one of the things that DuckDB did is they said, "Okay, we're gonna really, really solve this problem, and we're gonna make it so that you can just do SELECT * FROM CSV file name, and it will do the right thing." And they wrote a research paper on it. They put a full-time kind of PhD on it, and it keeps getting better, and they're writing more research papers on this. Versus if you look at other database companies, when I was at BigQuery, we put a college new grad on the CSV, and they worked on it for three months, and we're like, "Great, you shipped it. Go work on something else."

Jordan Tigani [49:43] And so there's all sorts of corner cases where things don't work, and then we basically say, "That's not our problem. Our problem is once you get into the database. It's your problem to go fix your CSV." And if you just think of it from the perspective of somebody who is trying to get work done, how much time do you spend wrangling broken CSV files? It's like, "Oh, this isn't working. This isn't working." And then so DuckDB can often make that just sort of magical and just sort of like, "Oh, it just works."

Jordan Tigani's entrepreneurial journey

Jordan Tigani [50:04] I don't have to think about it. And that's like a small example, but it's a way that, yes, the number of tools are maybe increasing, but hopefully we can still make life simpler.

Matt Turck [50:33] So to close, I'd love to go in a completely different direction and talk about your entrepreneurial experience. So you mentioned you hadn't started a company before, and you were most recently Chief Product Officer at SingleStore. What was the transition, for any technical person out there that may be thinking about starting their company? What was surprising in a bad way, but also in a good way?

Jordan Tigani [50:58] I always figured that there's people out there that are going to start companies, and then there's kind of normal people. And I was one of the normal people and never kind of sold computers out of my dorm room or had lawn-mowing businesses employing siblings. But I just sort of fell into this. And I think one of the lessons is, you might not think that you're an entrepreneur, but you too can do it. I think there's been a number of kind of surprises along the way, but I think one of them is that there's no one right way to do it.

Jordan Tigani [51:29] I kept expecting somebody who was brilliant, Lloyd, for example, to be able to tell me the answers. Like, how should I think about this? How should I do this? How much money should I raise? Who should I raise from? How should I hire my first engineers or my first salesperson? And you can listen to people who have been really successful.

Jordan Tigani [51:51] And that's one of the things that I love, is that founders tend to be so giving with their time to other founders, because otherwise I don't think I could have kind of figured anything out, because I had literally no idea what I was doing when I got started. But you listen to them and you're like, okay, well, this person did X and is telling me, you've got to do exactly what I did, of course.

Jordan Tigani [52:23] And then another person who's equally successful does exactly the opposite. And you're like, wait, well, how can these two things be? And you realize that, well, this was the right thing for that person, and this was the right thing for this other person. I've got to figure out what's right for my company and what's right for MotherDuck. It puts more pressure on your shoulders, but it also makes you realize that there's not one way.

Jordan Tigani [52:57] And I think one of the things that has been suggested is you've got to rely on your intuition and what feels right and what seems right to your company. I think an example from the early stages was I asked one founder, well, how much should I raise? And he said, raise as little as possible, because you can always raise more. If you raise from somebody good, they're not going to let you run out of money.

Jordan Tigani [53:24] And then I asked somebody else, and they said, raise as much as you can, because the only time you're ever at risk as a founder is when you're running out of money. If you control the board and you have plenty of money, you don't have to listen to anybody else. And both of those people are right, maybe more so and maybe less. I mean, if you raise too much money, there's certainly problems you can get into.

Jordan Tigani [53:57] If you raise too little money, there's certainly problems you can get into. And I think people—this was in 2022, kind of the tail end of the ZIRP era—and then people started finding, well, maybe you can't always just hold your hand out and get more. And so there were positives and negatives on both sides, and you've just got to figure out, okay, what's my decision framework and what works for me and what works for my company?

Matt Turck [54:17] What about the commercial side of things? Coming from a very technical background, how did you learn there? What was surprising? Again, with the caveat mentioned up front that you seem to be doing an extraordinary job at marketing and marketing positioning.

Jordan Tigani [54:43] So I was an engineer for 20 years, and then I kind of bounced back and forth between engineer and engineering manager. And I worked on a couple of projects that I thought were beautiful. I thought, wow, I'm just so proud of this, and they are dead, because the commercial value wasn't there, the company killed it, or the company's no longer there. And so I think one thing that taught me was, you've got to understand customers, you've got to understand the market.

Jordan Tigani [55:19] And so I was part of the team that helped start Google BigQuery, and at one point I ended up moving into product. And part of the reason that I did that is because, first of all, I was being asked to help hire the director of product for BigQuery, and I kept interviewing people. I'm like, well, this person just really doesn't get what's special about BigQuery, and this would be terrible if that person led the product. And then I thought, well, maybe I could do it, which felt totally weird to go from engineering to product.

Jordan Tigani [55:57] But that just opened my eyes in so many different ways because in engineering, you're designing with this sort of palette. You have these things that you can do, these data structures that you can move, these systems that you're building that connect in certain ways. On the product side, you have sort of the same ideas. You have customers and pricing and packaging and your marketing team and your go-to-market and all these things. You have to sort of design something coherent. And if you don't make it coherent, if you don't make it work, it doesn't work.

Jordan Tigani [56:29] Now, you don't get the same positive feedback every day that you do when you kind of submit code and it works and you see it running. On the other hand, I think you can build something more real and more enduring if you do take a product focus. So being a product manager for BigQuery did give me exposure to some of these other parts of the world that I don't think I would have had. And then I jumped into the deep end at SingleStore as the Chief Product Officer.

Jordan Tigani [57:00] And being in the C-suite, I got to be part of a lot of these discussions with the CRO and the CMO and the CEO and the CFO, and kind of hear how they were thinking about fundraising and sales comp plans and marketing. And I realized that marketing is war. When you're at Google, you think, oh, well, you don't need to do marketing. Who needs to do marketing? But then you realize outside of Google, nobody cares about what you're doing.

Jordan Tigani [57:31] And you have to basically win over their eyeballs. You have to convince them to care, and you have to give them a good reason to do so. Anyway, so I think it was a really good introduction to what it takes to see what a successful startup company—a much later-stage startup company, I mean, they were $100 million in ARR—what that looks like and all the pieces. And then you can sort of extrapolate back down to, okay, it does feel like a whole different thing when you're just starting a company and it's just you and a couple of people and don't even have a GitHub repository.

Jordan Tigani [58:08] And over time you can start to see, okay, well, this is the lineage that will get you to this thing towards the end, which then you look at actually successful public companies and you see, okay, this is what gets you there. So that was great. I highly recommend being part of a big tech company because you learn how to do things right.

Outro

Jordan Tigani [58:31] But then I think also, I recommend being part of a startup because you learn how to get things done quickly, and you also get exposure to sort of much more pieces of the puzzle than you would otherwise.

Matt Turck [58:37] Thank you so much for sharing the story and some lessons. Really appreciate your being here today.

Jordan Tigani [58:39] Thanks. Appreciated the conversation.

Matt Turck [59:01] Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already, and please leave us a positive review or comment on whichever platform you're watching or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.