Building The Database That Can Do It All | Tobie Morgan Hitchcock, CEO of SurrealDB
The MAD Podcast with Matt Turck · with Tobie Morgan Hitchcock, CEO, SurrealDB
Tobie Morgan Hitchcock is the CEO at SurrealDB. We cover why SurrealDB started as a way to replace four database backends, how its record IDs combine document, graph, and time-series queries without indexes, and why running models beside data avoids moving records through separate infrastructure.
Chapters
- 2:03 — What is SurrealDB?
- 2:53 — How did SurrealDB get started?
- 9:10 — The Challenges of Building a Database from Scratch
- 10:36 — Why SurrealDB Chose Rust
- 12:54 — A Deep Dive into SurrealDB’s Unique Features
- 19:30 — Why Now?
- 26:32 — What Sets SurrealDB Apart from Other Databases
- 30:01 — SurrealDB’s Role in the Future of AI and Machine Learning
- 32:45 — Why Developers Are Choosing SurrealDB
- 36:14 — What’s New in SurrealDB 2.0?
- 40:10 — SurrealDB Cloud: Scalability Meets Simplicity
- 42:21 — How SurrealDB Fits into the Competitive Database Landscape
- 45:37 — Early Lessons from Building SurrealDB
- 48:34 — Co-Founding SurrealDB with His Brother
Transcript
What is SurrealDB?
Matt Turck [1:15] Welcome, Tobie.
Tobie Morgan Hitchcock [1:17] Thank you for having me here.
Matt Turck [1:25] So this is actually one of the very first, if not the first, podcast that you do. Is that correct?
Tobie Morgan Hitchcock [1:33] This is the very first podcast I've ever been on. Done lots of streams, done video interviews, TV interviews, but never a podcast.
Matt Turck [2:03] Okay. All right. This is a world premiere. Very excited for it. Which is interesting because SurrealDB sort of popped seemingly out of nowhere a couple of years ago and got very noticeable and noticed success as an open-source project, really caught people's imagination. But while you have a great community and web presence, you haven't done too many of those things. So thank you for being here today. I really appreciate it. So maybe as a quick intro, what is the sort of 30-second elevator pitch for SurrealDB?
Tobie Morgan Hitchcock [2:18] So SurrealDB is reimagining databases. So effectively, the tool that people use to store data and then, once they've stored it, to query that data. The thing about SurrealDB is it makes it much simpler in terms of querying that data, and it means that you can reduce the number of databases that you have to work with as a developer, as an engineer, or as an organization by consolidating different database types together into one single platform that reduces cost, reduces development time, reduces the management of that application or that system as you go forward.
Tobie Morgan Hitchcock [2:51] So there's lots more to it, but in a nutshell, it's about simplification of the data backend.
How did SurrealDB get started?
Matt Turck [3:15] Okay, great. All right, so we're going to go into all of this, but maybe start with the story behind the company? I mentioned a minute ago that you guys sort of shot up seemingly out of nowhere in, I think, the fall of 2022 with a lot of open-source success. But is it fair to say it was one of those overnight success stories many years in the making?
Tobie Morgan Hitchcock [3:36] I'd love to say I wrote 300,000 lines of code on a weekend, but no, it didn't happen like that, unfortunately. So if you look at many databases out there, they're improving slightly on one particular thing. They may be improving on performance, or improving on what you can do with your data, or how it runs in the cloud, or something like that. But with SurrealDB, we had to reimagine many different parts of the database in order to make what SurrealDB is today.
Tobie Morgan Hitchcock [4:13] I guess where we start is why I decided to build SurrealDB in the first place, as opposed to how we started building it. So we were running traditional cloud applications on top of many different databases in a similar way to how other organizations are still doing it today. So, a time-series database, a document database, a graph database if you're doing analysis over related or connected data. And then we also had a relational database in the mix as well. Most organizations do.
Matt Turck [4:42] And to bring this to life, this is one of the many really fun parts of the SurrealDB story. So it was you and your brother Jamie, and you had started multiple applications to do different things, I think from sports analytics to other applications, and then you'd built a database underneath. Is that correct?
Tobie Morgan Hitchcock [5:00] Yeah, exactly. So we had a number of different products that we were building the technical side of. And in doing so, we wanted to simplify how we could do that so that we could run these applications efficiently without that much effort, as it were. And that was how SurrealDB came about. So we were using these different databases, and the management of that in terms of managing four different databases, even in AWS, ensuring that they're up, working as you want them to be, updated in the right ways when you're updating your application, that becomes time-consuming and complex, and we wanted to move away from that.
Tobie Morgan Hitchcock [5:46] And we wanted something simpler. At the time, there was Firebase, and Firebase enabled you to connect to a cloud instance and store your data and connect to it directly from the web browser and build your applications against it. But it was incredibly limited in terms of how you query your data. And today it's still limited in a similar way, that it's not really a database. It's what I would call a data platform or a data layer. We wanted to build a database that gave you the total flexibility to run your queries as you wanted to in your application, but with the same functionality that something like Firebase offered, but then also still act as a traditional database backend for those who needed that as well.
Tobie Morgan Hitchcock [6:25] It was actually funny. So my very first conference I spoke at was RustConf London. The entirety of SurrealDB is built in Rust, and we'll come to that a little bit later. But the first question we got from the audience was, look, it takes a long time to build a database. How did two of you build it? So it's actually only one of us. My brother is the marketing, the branding, the design, everything behind SurrealDB, but not the code.
Tobie Morgan Hitchcock [6:35] Yes. So it's just one person, and it was seven years in the making.
Matt Turck [6:59] So you all started in 2015 with this application and the backend. At what point did that sort of eureka moment—maybe that was not a moment—happen, that at some point you said, oh, well, actually the backend, that database we built to power those applications, is actually more interesting and has more potential than the applications themselves?
Tobie Morgan Hitchcock [7:09] So interestingly, I don't come from a database background at all. So I think if you were to have come from a database background, you probably wouldn't say, yep, we'll rewrite the entire database backend.
Matt Turck [7:10] Who'd be crazy enough to do that?
Tobie Morgan Hitchcock [7:38] Exactly. I think it was coming from a point of view of not actually understanding what was needed at that point. I'd had a lot of experience with lots of different database types, database internals, the codebase of many different databases, but I'd never really written my own. Definitely not. And I started with—so my master's thesis was in key-value stores and focusing on the performance of key-value stores when you're dealing with temporal data. So not just time-series data where you can order events by a timestamp, as it were, but actually temporal data, which enables you to go back in time and see data as it changes over time and gives you the ability to run entire queries over entire datasets as it existed at a particular time.
Tobie Morgan Hitchcock [8:20] So almost like a third dimension to your data. That was where a lot of my interest lay, and I did my master's thesis on that. But when you look at the power of what time brings to data, you can do so many things with it, but there's not a single database out there designed in that specific way to be able to do that. And that is where the idea for SurrealDB came from. And at that time, we wanted to reinvent this wheel.
Tobie Morgan Hitchcock [8:46] And many people say, well, why? Why do you need to? Why do you want to reinvent the wheel? Because if you want to add things like this temporal querying or the ability to combine multiple different types of data and querying into a single platform, you have to rethink how you do that. And so we started that in 2015. From a code perspective, I started that in 2015. And I think it was about two years later that we started using that in our first application.
Tobie Morgan Hitchcock [9:09] And I remember launching it on an e-commerce store, and we had a lot of visitors on that e-commerce store, and it just crashed instantly. But we went from there. And then eventually it was doing hundreds of thousands of queries effortlessly and on a single node. And we built from there.
The Challenges of Building a Database from Scratch
Matt Turck [9:22] Did you have any sense when you launched on that journey how sort of ambitious that was? Is it one of the things where, had you known, you might not have done it?
Tobie Morgan Hitchcock [9:47] Probably not. Definitely not. There's a blog post—I think it's a famous saying. There's also a blog post which was published, I think, in 2021 or something: "It takes 10 years to build a database." And when we actually launched SurrealDB as an open-source tool and it took off online, it had only been seven years in the making. And in the database world, you're not finished. You're never finished. You've always got to constantly improve and constantly build on top of that.
Tobie Morgan Hitchcock [10:03] But to see how users were enjoying and liking the aspects of what we'd built in SurrealDB, and then starting to use it in their organizations, that was kind of the turning point. So, seven years. But no, I think in the beginning, especially in the early days when you're running applications on it and you're trying to work out, at the same time, how to make the load work so that you can cater to the number of queries per second that are coming in, it's a great problem to have, of course.
Tobie Morgan Hitchcock [10:26] But it's also a massive learning curve. An enjoyable learning curve.
Matt Turck [10:32] Yeah. And when they said 10-year journey, I'm not sure they meant a 10-year journey by one guy.
Why SurrealDB Chose Rust
Tobie Morgan Hitchcock [10:36] Probably not. Probably not. I don't know what it actually refers to, but yeah.
Matt Turck [10:53] So maybe one last thing on the history of the company, then we can fast-forward to what it is today. But we started talking about Rust, and I think at some point there was a rewrite in Rust. Can you maybe walk us through what happened and what the decision was?
Tobie Morgan Hitchcock [11:11] So SerenityDB was entirely written in Golang, including our own data structures, which were custom data structures to deal with time in the key-value underlying storage engine. And then, in addition to that, we had the entire query engine, its own parser. Everything was handwritten. It was built in Golang. But we reached this point where we were having to—Golang's changed a lot in the last few years, so a lot of these may have been improved—but we were reaching the point where we were battling against the compiler itself and the language itself.
Tobie Morgan Hitchcock [11:46] We were building a lot of custom data serialization and data deserialization logic. As I say, the parser was written from scratch. We were still fighting against different aspects of the language to make it work according to how we wanted to make it work. And there's the idea that you build it, you get it out there, and then you improve from there. But we didn't want to release a tool which was working, but it wasn't working quite how we wanted it to work, and it wasn't working in the right direction to where we wanted to go.
Tobie Morgan Hitchcock [12:02] And so we made the decision to completely rewrite the entire codebase from Golang into Rust.
Matt Turck [12:04] Just to make things a little harder.
Tobie Morgan Hitchcock [12:22] Just to make it a little bit harder. At the time, I'd never written in Rust. I'd dabbled, I'd tried it, but it was new to me at the time. But it's very hard once you've launched an open-source product, especially one that you think, or you hope, is going to take off, to then go and completely change every aspect of it. What you'll see with other databases is they then design parts of it in Rust, and then gradually more and more of that becomes Rust over time, and maybe some of it doesn't.
Tobie Morgan Hitchcock [12:52] But it's much easier when you go forward from the basis of everything being built in Rust. We can now build a team of people around the Rust code, and we can go forward in a much easier way. So that was the idea. And it also benefited us because the Rust community is amazing, and you've got some incredible engineers out there. And that helped us when we launched.
A Deep Dive into SurrealDB’s Unique Features
Matt Turck [13:31] In addition, launch was a couple of years ago. Fast forward to today, you have 27,000 GitHub stars, which is one of the key metrics, and one can argue what those mean, but directionally very impressive. You have hundreds of thousands of downloads. Also, last week, and we're going to talk about this, you're about, or in the process, I guess, as this podcast gets published, to launch the cloud version, which is going to be available. So we're going to talk about all of this, but let's get into the product itself.
Matt Turck [14:07] So SurrealDB is billed as the ultimate multi-model database. And not to be too cute about it, but it's model and it's not modal. So it's M-O-D-E-L, not M-O-D-A-L, which is what you see in a lot of generative AI conversation. Do you want to explain what that is and then maybe walk us through the different parts of the product that are noteworthy?
Tobie Morgan Hitchcock [14:19] Yeah, you'll be interested, like, the number of people who interchange those words. And I guess you can, in a way. It is multi-model and it is multimodal, but—
Matt Turck [14:20] Oh, that's how you pronounce it. All right.
Tobie Morgan Hitchcock [14:47] Well, I do. I'm not sure it's all right. So multimodal is about dealing with different modes, so images, audio, video. Model, I guess, is a different way of saying it, but it's looking at different ways of, in the database sense at least, it's looking at different ways of querying and storing different types of data. So in SurrealDB, you can store traditional document data, and it is a document database under the hood. So it's storing records that look a bit like JSON, in a way.
Tobie Morgan Hitchcock [15:15] You can then improve and augment that data using graph techniques. So we'll come back to what graph is in a second. And then time-series data, so effectively a stream of events tagged with a timestamp, as it were. I think the interesting thing that people find when they come to SurrealDB, because you have graph databases out there and you have document databases and you have time-series databases. Just to add to that, you can also query things in a key-value-like way.
Tobie Morgan Hitchcock [15:41] So if you want to get a single record, it's very efficient in just getting a single record, which is great in certain industries where they want speed and performance. And in addition to that, it's got an SQL-like query language. So it's very similar to traditional SQL that you might find in a relational database, but it doesn't have joins. And so when people come to SurrealDB, they usually come from the perspective of a relational database, which is tabular, or a document database, which is tabular, albeit with nested fields.
Tobie Morgan Hitchcock [16:21] And they usually have to relearn how they insert their data and manage their data in those different types of databases. What they find when they come into SurrealDB is you can start with your data in that way, but then you can augment that with relationships. And actually, relationships and graph is really the way that humans think about the world. So we think about the world in terms of types: people, animals, orders, items, products. But then we augment that with relationships.
Tobie Morgan Hitchcock [16:51] So I know you, you bought an item. And when you think about it in that way, modelling your data in your application or in your database, when you can deal with traditional tables and collections of data, but then you can bring relationships into it, makes a lot of sense. So bringing these different models together in a single platform, but not forcing you to have to use one model or the other. You can use all of them at the same time. You can start with one, you can build on that.
Tobie Morgan Hitchcock [17:11] You can start without any structured schema. You can improve and add structured schema as you go along with your application. It has a lot of flexibility to it, and it makes sense to a lot of engineers, regardless of which database type or which experience they've come from before.
Matt Turck [17:21] All right, so it can do graph, it can do document, it can do transactional. It can also do a lot of this in real time.
Tobie Morgan Hitchcock [17:46] Yeah. It's a transactional database, so it is designed for reads and writes, as opposed to an analytical database, which is traditionally more read-heavy over large amounts of data. But there are two core things in SurrealDB that set it kind of in the middle as a hybrid transactional analytical database. And that is, the first thing is the storage layer of that data is separated from the compute layer. So SurrealDB can run as an embedded database like SQLite, it can run as a single node like Postgres, or it can run as a distributed cluster.
Tobie Morgan Hitchcock [18:16] And you can scale this storage layer and you can scale the compute layer. What that means is if you've got large amounts of data, you can scale that independently. And then if you want to have lots of reads or you want to be doing lots of analytical queries, you can scale the compute layer out and then scale it back down as you need to. Or vice versa, depending on the needs of whether you've got lots of reads in your application, lots of writes in your application, and so on.
Tobie Morgan Hitchcock [18:48] The second thing is when you combine traditional row-based data or document-based data with graph and you model it in certain ways, you can actually perform very powerful real-time analytics. So the word real-time has many different meanings in different technical environments, but effectively you're able to perform an analytical query in an expected amount of time. So instead of having to run this query and it takes a certain number of hours and you can process something for the next day, you can actually run queries and get back results very, very quickly.
Tobie Morgan Hitchcock [19:29] So combining the transactional approach, which is when you insert data in the database, you know it exists there and it exists for everybody else reading that data after that point, but also when you want to run analytical queries using key-value or time series or graph or traditional, just row-based tabular data, it's very powerful. It enables you to mix and match those different types of queries that you might need in an application or in your backend, as it were.
Why Now?
Matt Turck [20:09] So my mental model for what you guys are doing from a market analysis perspective is that we've just gone through a period of, I don't know, 10 or 15 years of unbundling of the database market. So you started with the Oracles of the world, and then as a next phase, you had all the NoSQL databases, the MongoDBs, and then you had the time series databases and the graph databases, as we discussed. But as separate best-of-breed kind of products. And it seems that you are pioneering and leading the charge in terms of rebundling of the database market.
Matt Turck [20:50] And I'm curious about the why now of that from a market perspective, but also from a technical perspective. Why is it possible to do this in the first place? You would hear a lot in the market that you cannot be the database that does it all. You have to choose and make trade-offs. So why is this happening? Why are you guys able to do this today?
Tobie Morgan Hitchcock [21:15] Yeah, so the separation of storage from compute is an important part in that. So being able to scale those out dependent on what you need in your application. The biggest change in SurrealDB is how we deal with the IDs of documents. So in Postgres or a relational database, something like MySQL, you traditionally have an ID field, which is an integer. So 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, and so on. And you can change that, but that's pretty much how it's initially formed.
Tobie Morgan Hitchcock [21:44] In something like MongoDB, you have an object ID, and that is a string, for instance. And then graph databases are slightly different. But in SurrealDB, the ID is made up of the table name that the record belongs in and the ID itself. So that ID can be a number, it can be a string, it can be a UUID, and it can be something more complex, which I'll come to in a minute. By changing that one small piece of technology enables you to bring all these models together.
Tobie Morgan Hitchcock [22:15] So if I want to do very quick requests to get just one record from the data store, maybe I'm in ads and I need very quick response time for a very particular record, I can go and get person Toby, and I don't have to use an index. I don't have to know anything else about the data. I can just go and pull that record immediately from the store without having to access anything else in the table. When you use that inside, that ID can then be pushed into documents as well.
Tobie Morgan Hitchcock [22:42] So when you use that inside other types of queries, you can then say, look, I have an orders collection inside this record, and the orders has 3 IDs. They point to the order table, and therefore I want to find out which orders this user has made. I can now fetch user Toby, and I can fetch his orders, and you're now going to get 4 records. That's graph. You can then go a step beyond that and you can say person Toby ordered product laptop, and then you can traverse bidirectionally between those 2 different types of records.
Tobie Morgan Hitchcock [23:25] If you look at a lot of organisations, they are actually storing a lot of these. They're either specializing in different database platforms. They've got time series databases, they've got graph databases, and they've got traditional document and relational. And often you have to have people expert in those areas, in those different databases. And then you have to make sure that your application, your middleware, or API is connecting to these different database platforms and bringing that all together. And people often look at the performance of one database and they say, well, this is really quick for document data, so you're not going to be able to compare.
Tobie Morgan Hitchcock [24:04] But they don't look at the overall platform, as it were, the overall architecture. And when you're having to combine time series data with graph, with document for a particular dashboard or an API, then you have to start taking into consideration all of the performance of all of the platforms put together. Coming back to what I was saying before, a lot of people say you can't get the performance when you're dealing with lots of different models. But because of this record ID that we have in SurrealDB, we're able to have good performance from a key-value perspective, have good performance from a document perspective, bring graph into that and have good performance in the same way on graph.
Tobie Morgan Hitchcock [24:45] And then for time series, we go one step further. So the ID, which I said before could be an integer, a string, so a number, a string, or a UUID, can actually also be an array of values. So now I want to create a temperature table, and the temperature table has—each record has its ID, has an array, and that array is—the first one is the city, London, and the second one is the date. And then very efficiently, I can query all temperatures in London between date A and date B without having to use indexes, without having to use any table scans or anything like that.
Tobie Morgan Hitchcock [25:24] So it's very efficient. So you can combine these different models together. Importantly, there's always a use case for a time series database. There's always a use case for a graph database as well. So graph thinks about data usually in a different way. You usually think about data in terms of metadata, which describes objects, and then you filter based on that metadata. In SurrealDB, we still have collections or tables. So you could create one big object table if you wanted to, and just filter and describe data based on the fields in that table.
Tobie Morgan Hitchcock [25:59] But it's not really the way that you would come to SurrealDB and do that. And time series as well. Usually time series databases are specialists in evicting data and the granularity of data as you go on in time if you don't need that data for your dashboards or anything like that. So there are different requirements, but a lot of organizations are working on analysis over time-based data with graph, with document. And actually, if you're doing that with 3 or 4 different database backends, that can become incredibly complicated.
Tobie Morgan Hitchcock [26:16] The performance can degrade over time as well, or just be harder to reach because you're working with different platforms. And that's where SurrealDB shines.
What Sets SurrealDB Apart from Other Databases
Matt Turck [26:59] We'll double-click on this in a second because that's really interesting. But before we do that, what is the claim then in terms of performance? Are you saying, hey, we can be good enough across those different dimensions, therefore use us? Or is part of the claim that, hey, compared to each best of breed, we're actually pretty close or almost there? I guess I'm going to sort of trade off. It sounds amazing, but where's the catch?
Tobie Morgan Hitchcock [27:21] Yeah. So for us, we know that in certain areas, in certain models, we can actually beat the competition out there, even whilst having all these different models in our platform. As I say, some of the functionality of what you might see in a time-series database or a graph database might not necessarily be there, and it's not really a priority of SurrealDB to do that. But in terms of performance, that's definitely a key part of what SurrealDB is aiming to achieve.
Tobie Morgan Hitchcock [27:46] In addition to that, if you're then consolidating different databases together and you're comparing that to a traditional backend which has multiple platforms, then definitely, the aim is to be quicker. Having said that, there are always going to be scenarios where another database might make sense. If you're dealing with graph data that all sits inside memory and it's being stored in a graph data structure, then that's always going to be quicker on a graph database that is sitting on a single node.
Tobie Morgan Hitchcock [28:23] For SurrealDB, everything that we build has to work in an embedded environment, a single-node environment, so a single server. But then we have the ability to scale up. So if you need to have more writes, more reads, you can scale up the compute, you can scale up the storage. So we have that ability as well, which is not necessarily true of other databases, or scaling is sometimes a difficulty with other databases.
Matt Turck [28:33] And then you and I were chatting right before this. You're going to have performance benchmarks coming up in the next few months.
Tobie Morgan Hitchcock [28:59] It's always a hard one. As a database, it's always a contentious issue, isn't it? As a database, you've got to put out performance benchmarks, and we're working on those internally. We're going to be publishing them soon. But you've got to compare yourself with, in our case, many different database platforms. So people want to know how we compare with the relational databases, the graph databases, how we compare with databases at scale, how we compare with databases that are embedded. Because it's so flexible, you've got to be able to compare yourself with so many different types.
Tobie Morgan Hitchcock [29:30] And then also you've got to test yourself in the right way, benchmark yourself in the right way. So for us, it's about supporting many of the open-source benchmarks that are already out there and building our results into those frameworks, but also comparing ourselves with other databases in the right way. The benchmark game in the database world is a matter of fine-tuning and perfecting what works for you and how it presents well for you. But I don't think that's right. SurrealDB is about simplifying developers' lives, simplifying organisations' lives, and making it better and quicker to get to the end goal of building your application or building the platform that you need.
Tobie Morgan Hitchcock [29:51] And to do that, it needs to be representative of many of the different environments, or as many as possible of the different environments that people are going to be running it in.
SurrealDB’s Role in the Future of AI and Machine Learning
Matt Turck [30:11] Still on the product front, obviously the trendy topic du jour is AI, and you guys do some very interesting things on that front. Can you just double-click on it, how you think about SurrealDB as part of the emerging AI stack?
Tobie Morgan Hitchcock [30:34] Yeah. So if you take everything I've said about consolidation and you then look at the world of AI and machine learning pipelines, it gets even more complex and confusing, as it were. For us, when you're working with machine learning models or artificial intelligence, whichever way you look at it, you're dealing with data, and your data is residing in a database. So for us, one of the pieces of functionality we have in SurrealDB is the ability to bring a model right inside the database so that you can run that model.
Tobie Morgan Hitchcock [31:09] It's a custom-trained model or an off-the-shelf model. You can run that model right alongside your data if you want to run it on all the records in a table. Let's say you want to update the predicted house price of all the houses in a table. You can run it very efficiently. You don't have to push that data out into Kubernetes clusters through different platforms, wait for asynchronous events to come back. And having finished, you can just push that model into the database and have it reside next to the data.
Tobie Morgan Hitchcock [31:46] So you're not having to deal with a lot of the administrative overhead of dealing with microservices and Kubernetes clusters and containers and everything like that. And it comes back to the question you asked earlier, which is: why rebundle? Everyone built microservice-based architectures because that was the easiest thing to do. It meant that you could keep components of your code small and you could build up your application from there. But it also led to a large, massive amount of complexity. And ultimately, it all comes back to the data.
Tobie Morgan Hitchcock [32:11] If you look at the biggest organisations in the world, they're leveraging that data in better and bigger ways. And for SurrealDB, for us, we want to make it easier to work with, to query, to learn from your data. So being able to do that in a single place without having to think about the infrastructure around that is really what we're trying to do.
Matt Turck [32:17] So there's a concept of bringing the models to the data rather than the data to the models. Is that fair?
Tobie Morgan Hitchcock [32:37] Exactly. And underneath the hood, this is still powered by GPUs. It's still pushing through to the same technology that traditional models or models out in other platforms are still running, but it enables you to bring that closer to the data. So you don't have to think about the infrastructure around running that processing or understanding the models against your data. You're just having to think about how does the model work, what is it giving me, what is it resulting in, and then how do I run that on my data?
Why Developers Are Choosing SurrealDB
Matt Turck [33:12] All right, so switching maybe from product into sort of like customers and use cases, and maybe to drive home a couple of things that you've alluded to in the earlier part of the discussion. So if I'm a customer, it seems that I have a couple of reasons to work with SurrealDB. One of them—is that fair?—is just simplicity, right? Is that just one database instead of multiple? Is that cheaper? Is it simpler? Maybe talk to that.
Tobie Morgan Hitchcock [33:40] I think cost definitely comes into it, but I don't think it's the main reason why people are turning to SurrealDB. I think longer term, people see the gains of being able to consolidate different databases. But especially if you look at the large organizations using us, they're not coming to us and saying, "I'm going to replace four different databases that I've spent the last 10 years implementing." It's usually being able to solve one problem to begin with. And that problem is usually that they're having to work with many different data types, or they're not able to work with the data types that they want to, or not able to query in an efficient way.
Tobie Morgan Hitchcock [34:15] And so they start using SurrealDB, and then they can see the benefit of bringing more data types and more models and more ways of querying to SurrealDB over time. That's definitely what we're seeing. We're seeing use cases from real-time analytics. So the combination of document and graph to bring kind of expected query times to things that need to happen in real time, as it were. We're seeing use cases where it's purely being used as a key-value store the majority of the time, but then they want to expand upon that and have filtering and document data storage.
Tobie Morgan Hitchcock [34:39] So it really is about being able to bring those models together. And I think in the long run, that does reduce costs, it does reduce development time, and it reduces the complexity of people's infrastructure.
Matt Turck [35:03] Yeah. So reducing complexity and, as you just said, enabling specific use cases. And to double-click on, again, I think what you just said, you were talking before this, before we started recording, about a customer using SurrealDB for product search, a combination of vector search and graph. Is that correct?
Tobie Morgan Hitchcock [35:03] Yeah.
Matt Turck [35:09] So a mixture of document, storing those records, graph, and product search on an e-commerce website.
Tobie Morgan Hitchcock [35:39] Exactly. So, e-commerce website, graph to analyze the related connectivity of that data. And specifically, it's product recommendation. So being able to detect in real time, not based on analytics of yesterday's data, but actually in real time on the events that are happening on the website right now, how do we analyze which products to recommend to this user based on the behaviors of other users, based on the behavior of this user, and based on many other attributes. So being able to combine graph with the document—the document is what stores the actual data, the graph is what relates that data—and then being able to do that at scale.
Tobie Morgan Hitchcock [36:13] And I think another key part of being able to scale out your database is that gives you the availability aspect. So being able to say, look, I can spin up 15 nodes or 10 nodes, but when three of them go down, my database is still up and running, as opposed to when one node goes down, I now have an issue. So that obviously comes with the ability to scale out your compute and scale up your storage.
What’s New in SurrealDB 2.0?
Matt Turck [36:34] People can go on the website and look at all the details, and you have great videos and all the things. But what are some highlights? In particular, I think you did a lot of work around security. What would you highlight?
Tobie Morgan Hitchcock [37:03] Condensing one year of work into 30 minutes was hard enough. Let's condense it into one minute. One of the reasons that people love working with SurrealDB is the query language. It's SQL-like, so it makes sense to people who are coming from a relational database, but it brings these concepts of graph and key-value into it. For us, extending and building upon that language is important. So, we've built in new functionality that enables people to work in a more programmatic way.
Tobie Morgan Hitchcock [37:38] In addition to that, security and stability. We're built in Rust, but being able to ensure that users authenticate to the database correctly. One feature we haven't spoken about is being able to lock down your database. So, being able to say, look, certain users can only see these records, or they can only run these functions, or they can only see these particular fields. One of the other features we just released in beta is GraphQL support. So, being able to query your database.
Tobie Morgan Hitchcock [38:07] This is not a middleware layer that sits in front. This is built right into the database. When you define your tables and you add your data to the tables and you define fields on that data as well, your GraphQL is automatically generated. You can configure how it is output, but it's automatically generated and you can query that immediately. So there are a lot of organizations out there running and building on GraphQL. And so that now is an instant kind of configuration flag when you get going with your database.
Tobie Morgan Hitchcock [38:19] So, temporal querying. We've gone full circle. This is how it started in 2015.
Matt Turck [38:45] Yeah.
Tobie Morgan Hitchcock [39:11] But actually implementing this in a way that's scalable and performant is just really hard. We will be the first database that enables you to time travel in your data in the different types of models that you can work with in SurrealDB. So now, when you're using the embedded storage engine called SurrealKV, you can add a version clause at the end of your SELECT query or INSERT query, and you can actually time travel back, which isn't that big a thing. You can do this in some other databases, and you can do it if you have a time column and you filter on that time column.
Tobie Morgan Hitchcock [39:45] The difference is in SurrealDB, you can take your document database or document data that relates to other data through graph edges. And you can also take your time series data and you can go back in time. So you could say, I want to find all products that this person has purchased, and I want to find other people who have purchased that product and what products they purchased. So that's the graph. And I want to do that as of now. But it's now come up to December, and I want to know for the holiday season what products to recommend based on last year's activity.
Tobie Morgan Hitchcock [40:09] Now I can time travel back in time a year, see what the data looked like a year ago, and effectively, immediately, I'm able to look at that data, look at how the entirety of my database—from graph, from document, from key-value, from time series—looked at a particular point in time.
SurrealDB Cloud: Scalability Meets Simplicity
Matt Turck [40:25] And by the time we publish this in a few days, you'll be about to open your cloud product from private beta to general availability.
Tobie Morgan Hitchcock [40:27] Absolutely.
Matt Turck [40:33] Any word on that? How does that work? What should people know, and how do they sign up?
Tobie Morgan Hitchcock [40:57] Yeah, so we wanted to make it a Surreal experience, right? So effectively, there's not really much difference. You can run Surreal locally on your own laptops, in your own clouds, or you can run it in the cloud. And that experience is supposed to be very similar. It's supposed to be simple, quick to get going with. The difference in Surreal Cloud is that it's all supposed to be super easy. But in addition to that, we don't just separate storage from compute.
Tobie Morgan Hitchcock [41:26] We're actually going a step further than that, and we're separating storage from compute. So how this works is, we have a storage engine running in the cloud for all of the clusters that are spun up by users. We then connect our compute loads to that, but the data itself is actually residing in object storage, so S3 in an AWS situation, and that gives us nine nines of durability on the data that resides there. It enables us to scale out as well.
Tobie Morgan Hitchcock [41:51] And at the same time, the storage layer in front is acting as a cache. So you get the same performance that you'd have in SurrealDB traditionally. You're able to query in the same way transactionally. So everything that you write to the database, when it's written and confirmed, every other reader sees that data. But then you're safe in the knowledge that your data is residing on S3. And if your instances go down, then they can be brought up again as well.
Tobie Morgan Hitchcock [42:16] And this is really important in analytical use cases. So being able to spin up large clusters that read large amounts of data, and then when you don't need those clusters, if you want to scale them down, that can now be possible. And also what we'd be working towards is the ability to autoscale up and down based on the needs of your application.
How SurrealDB Fits into the Competitive Database Landscape
Matt Turck [42:55] So maybe zooming out, I would be very curious about your thoughts on the database market in general, whether you look at it from a competitive angle or not. But it's arguably both the largest market in enterprise software and also the most competitive. And not to pick on them necessarily, but there's this whole group of Postgres-based modern databases, some of which are doing great things with great traction. What do you make of it, and where do you think it all goes?
Tobie Morgan Hitchcock [43:04] I just want to start by saying, look, I love data, I love databases.
Matt Turck [43:06] We're all friends.
Tobie Morgan Hitchcock [43:34] We're all friends. I wouldn't be here if I hadn't had all of those experiences and those learnings on the different databases that I was running. And obviously, there are some great companies out there pushing incredible technology. I think if you're looking specifically at SurrealDB, the reason we are able to have so much interest from developers and engineers is because we're not doing things in the same way that it's been done before. We're not building on top of Postgres and therefore limited by what Postgres has to offer.
Tobie Morgan Hitchcock [44:09] We were able to build out our core infrastructure and the core engine of SurrealDB, and that gives us additional benefits down the line. It enables us to support different models, it enables us to scale or run embedded, it enables us to run in so many different ways. There's a reason that SurrealDB has seen so much interest from users and massive organizations around the world, and that is because there is a gap in what the database market is offering. A lot of people, as I said at the very beginning, a lot of organizations are focused on improving the performance of an analytical database or reducing the cost of an analytical database or adding some functionality to an alternative product that's out there already, or offering a cloud version of an open-source product.
Tobie Morgan Hitchcock [44:46] But no one's really come along and said, how can we completely reimagine this and take concepts of SQL, but not fully lift SQL from the standards? Take concepts of graph, but not be a graph database. Take concepts of time series, but not be a time series database. And we're looking at it from the point of view of how do we really help developers and organizations improve their developer experience, improve their application build times, their application development times, and improve the management process once launched.
Tobie Morgan Hitchcock [45:17] Rather than looking at milliseconds or nanoseconds or cost. And that's, as I say, the database world is incredible. There's incredible products out there, but there's a reason that SurrealDB has seen so much interest.
Early Lessons from Building SurrealDB
Matt Turck [45:52] You've had a rapid and, I imagine, pretty dramatic transition from being a team of two, or, as we discussed, a team of one, building a very technical product to now running a thriving young startup. Anything that surprised you, or early lessons in particular for any very technical founder listening to this? Anything that comes to mind in terms of your early journey in that phase?
Tobie Morgan Hitchcock [46:23] I think there are learnings as a business, as you present yourself externally, and there are learnings internally as well. One of the hardest things for me—I built this code in the very beginning. I spent seven years on it, and bringing people into that codebase is tough. Still to this day, I'm overseeing a lot of what is going into the core engine. And bit by bit, that's reducing. But it's almost like giving a baby away. As a highly technical product, like a database, specifically like SurrealDB, the outside world can love you and hate you at the same time.
Tobie Morgan Hitchcock [46:58] They can love what you're building, but at the same time, they expect you to be where a 20-year-old database is in terms of functionality or performance. And I think sometimes people forget that this is an incredibly complex product to build, and it takes time as well. And so, one of the challenges is: how do you focus on engineering? You have to devote time and resources and engineering talent to building this incredibly complex source code and an incredibly complex product whilst, at the same time, catering to the community and their needs and what they want and what they want to see, even if that might potentially slow down what you're building.
Tobie Morgan Hitchcock [47:24] So it's a balance. It's definitely a fun journey.
Matt Turck [47:25] Yeah.
Tobie Morgan Hitchcock [47:53] We have had the benefit of being able to hire from the community, where you can already, up front, see the quality and the passion of people before you bring them into your organization. And at the same time, SurrealDB came about because I was able to spend time thinking about how to change a particular market, not by building according to a specification. And so the engineers that we have and we want to have in SurrealDB need to also have that imagination and that freedom to create, because I think it's through that freedom that you get amazing ideas, features, and functionality.
Tobie Morgan Hitchcock [48:33] And I think it's apparent that organizations, users, and developers love what we're building. But it's important that we keep having that flexibility in the engineering team in order to let the engineers work on really what they're passionate about, in order to deliver, longer term, what we want to deliver.
Co-Founding SurrealDB with His Brother
Matt Turck [48:56] Your co-founder is your brother, as we mentioned upfront, which is a reasonably rare thing. There are married couples. I mean, famously, there's the Collison brothers. So it's been done before very successfully. But what are your thoughts on, I don't know, the pros and cons of the experience of working with a very close sibling as a co-founder?
Tobie Morgan Hitchcock [49:23] I mean, you've met my brother. For anybody who has met my brother, we are very different. No one would actually guess that we are related at all. He is non-technical, as it were. He can build websites and code. He does code our website, but he is non-technical. His expertise, his passion, lies in the branding, the design, the external perception, the external imagery of the business. Anyone would say that normally wouldn't work, but I think there are two things.
Tobie Morgan Hitchcock [49:57] First of all, we have the technological understanding and the branding understanding. And I think that's quite rare, especially at an early stage, for founders to have. The other thing is we have trust. I absolutely trust—not necessarily trust about what my brother's going to do, but trust that we have the same vision, and trust that we have the same understanding of where we want to go and how we want to do it. And that makes running your business easier if you don't have to constantly be thinking about your co-founders and everything else beyond that.
Tobie Morgan Hitchcock [50:34] Because a lot of time goes into engineering and building the organisation and talking to users and organisations who are using us. And there's not much time to worry about the politics of it. And I don't know the statistic; you're probably better than me on this, but a lot of startups fail because of the falling out of founders. I'm not sure what that number is, but it's high. And so having that trust and having the mutual understanding of where you're going just makes it so much easier.
Matt Turck [50:53] And presumably you have a sort of built-in conflict resolution, or at least a long history, as any siblings, and particularly brothers, I guess.
Tobie Morgan Hitchcock [51:18] Twenty years ago, it would have been a fight where I would have lost, and then that would have been the conflict resolution. I think we're often on the same page, to be honest. He doesn't direct the direction of the product or have any influence over the direction of the product because that is my knowledge. But vice versa, I don't direct—or have any influence over—the direction of the brand or the marketing side. That is where his passion lies and where his expertise lies.
Tobie Morgan Hitchcock [51:44] And so there's no crossover in those areas. But I think also it's important to note in any organisation, but also importantly in a startup as well, you can't just focus on the brand or you can't just focus on the tech. You have to focus on both. For SurrealDB to be successful, we have to be an incredible technological product, and we have to have a good brand to go with it. We're talking to large organisations, and the brand and the business matter as much as the technology.
Tobie Morgan Hitchcock [52:03] And you talk to developers and engineers who are working on side projects, the technology matters more than the brand, but you have to balance both of those out if you want to build the company that we want to build.
Matt Turck [52:28] Well, and that certainly shines through in everything that you guys do, in terms of whether the website or the videos or the presentations or all the things. There's clearly a brand and a style to it that is very noticeable. Okay, well, wonderful. That feels like a wonderful place to leave it. Thank you very much for doing this. We're glad to be your first podcast and look forward to seeing the cloud product out and all the great things that you guys will do.
Tobie Morgan Hitchcock [52:43] Awesome. Yeah, I've got some exciting stuff ahead. So thank you for having me, and great first podcast.
Matt Turck [53:04] Thanks, Tobie. Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.