Chroma: Programmable AI-Native Databases with Co-Founder Jeff Huber

The MAD Podcast with Matt Turck · with Jeff Huber, Co-Founder, Chroma

Jeff Huber is the Co-Founder at Chroma. We cover why programmable memory lets developers deterministically specify the knowledge and tools an LLM uses, why grounding a model with retrieved documents is the state-of-the-art response to hallucinations, and why AI-native databases need both transactional and analytical workloads.

Watch on YouTube

Chapters

  1. 0:00 — Full episode

Transcript

Full episode

Matt Turck [0:40] Welcome, Jeff.

Jeff Huber [0:41] Thank you.

Matt Turck [1:14] You are the CEO of Chroma. Chroma has raised about $20 million in venture funding, most recently an $18 million seed round, a large seed round led by Quiet Capital with a bunch of famous people from the Valley and beyond as angel investors. So congratulations on all the success so far. Would love to maybe start with your background. What led you to start Chroma, and what were you doing before that?

Jeff Huber [1:44] Yeah, for sure. Most previous to Chroma, I worked on a YC-backed company for about seven years. Long journey, lots of stories I could share. I think the one thing that I learned most poignantly during that time was the difficulty and pain required to turn a machine learning model from demo to production. It's just so hard to get any level of reliability out of these things. Any level of nines, even 90%, can be quite difficult in a lot of cases. So this was a shared pain with my co-founder Anton as well.

Jeff Huber [2:00] He had worked on ML at Nuro and Facebook. And we thought that there was a better way. We wanted to make the tool that we wanted to have when we were developing machine learning applications. And so we started down this exploration of trying to understand how interpretability of latent space by looking at embeddings—we can get into some more definitions here that might be helpful—by looking at embeddings, doing embedding search, doing analytics over embedding space, you could give developers at the minimum a divining rod, if not a compass, to be able to improve their models and get to the level of reliability they want to have in their applications.

Matt Turck [2:33] Great. And so when did you start the company?

Jeff Huber [2:37] We did the pre-seed round in May, April of last year.

Matt Turck [2:39] Oh, last year. Okay. So it's still very new.

Jeff Huber [2:42] Okay. Yes, we've been around 14 months or something like that. Yeah.

Matt Turck [2:50] You've done a terrific job bursting onto the scene because you don't come across as a 14-month-old company. So well done on that.

Jeff Huber [2:50] Yeah.

Matt Turck [3:01] All right, so let's jump into it. What is Chroma, first at a high level? And then we'll get into some definitions, as you suggested, around embeddings maybe.

Jeff Huber [3:29] Yeah. Obviously, what Chroma is today, what Chroma is going to be in the fullness of time, all of these things are interesting. I'll talk about today because that's the most honest, that might land the best. Chroma is an open-source vector database. Our goal is to give developers the tools to create programmable memory for language models. Language models are obviously incredibly powerful, incredibly useful, but difficult to tame. The Shoggoth meme comes to mind for those of you that are terminally online like myself.

Jeff Huber [4:04] The green monster with all the tentacles and eyes all over it. And how do we tame this beast? How do we tame this beast and make it useful and reliable? And Chroma's belief is that programmable memory, so developers being able to set deterministically, hey, language model, this is the knowledge you should know about. This is the knowledge you should use. These are the tools you should know about. These are the tools you should use, how to use them, poses a path to, again, bringing language models into every use case, being able to depend on them, rely on them.

Jeff Huber [4:15] I won't say make them aligned, but maybe that too. Okay, great.

Matt Turck [4:29] So we're still early as a community learning about all of the things. So a vector database enables one to store and process embeddings. What is an embedding, and why does it matter?

Jeff Huber [4:56] Yeah, for sure. So an embedding is just a point in higher-dimensional space. So you can imagine a map as latitude and longitude. You place a pin on the map, you have two numbers, latitude and longitude, and that gives you some semantic idea of what that point means. It's on the continent of North America. It's near other cities. You can cluster, and you can understand that point's context given its neighbors. Embedding spaces are not two-dimensional. They're hundreds of dimensions or thousands of dimensions or even tens of thousands of dimensions.

Jeff Huber [5:25] And they're usually trained—the models that generate these embeddings, embeddings is a fancy word for a list of numbers, it's literally just a list of numbers that represents a point—are usually trained what's called contrastively learned for semantic similarity. So the idea is that like statements will be near each other in embedding space. You already saw some of this stuff from our earlier demos. So a good example of this that I always use is if you're trying to search an HR knowledge base for what your company's time-off policy is.

Jeff Huber [6:02] Maybe you're using the wrong language. Maybe you're saying, hey, what's my time-off policy? But you also want to get back, if you're European, the company's holiday policy. If you're an American, the company's vacation policy. All three of those things are semantically similar. A good embedding model would retrieve those sections of the HR knowledge base and bring them back either to you, the person who searched it, or to the language model for further synthesis. That ability to do fuzzy matching is quite helpful, or fuzzy search is quite useful and quite powerful over conventional text search, which is, you'll have a lot of misses.

Matt Turck [6:20] At an even higher level, you need to turn all of this into numbers because that's what machine learning models want.

Jeff Huber [6:43] Correct. Yeah. At a technical level, one way to understand it is you're taking unstructured data, text, images, et cetera, and you are vectorizing it. You're turning it into numbers that mean something, that have semantic meaning. And that allows you to do things like, at prompt time, when a user types a prompt into a box or other, be able to figure out at prompt time what is the relevant information we need right now, and then pull that into the context window and kind of do this.

Jeff Huber [7:07] I like to use the analogy from The Matrix. When Tank downloads the information in Neo's brain for how to do kung fu, and then he has the famous line, "I know kung fu," and he goes and fights Morpheus and it's awesome. That idea of being able to download instructions into the brain of something so that it can then go do something useful, I think, is quite analogous. Obviously, it's an abstraction, but it's quite analogous to the power that a vector database can bring to a language model.

Matt Turck [7:31] And the step that's upstream from you guys, and I don't think that you do that because you focus on the storage part, but the conversion, how is that typically done? Are there names that people should know in terms of models and companies that do the conversion?

Jeff Huber [7:57] Yeah, exactly. So there are language models. I'm sure many of you, all of you hopefully, have used ChatGPT and you're pretty familiar with what that is. There are also something called embedding models. The output of embedding models is not a paragraph of text that goes into a chatbot. The output of an embedding model is this list of numbers. So an embedding model is still a machine learning model. It's still trained. OpenAI has one called Ada-2 that just got, today, 75% cheaper.

Jeff Huber [8:14] There are also closed-source embedding models from Cohere. Google PaLM has one. And then there is a plethora of open-source embedding models as well that many people in the community find to be very good, actually.

Matt Turck [8:26] So you're an enterprise user, you've got a bunch of text content or audio or video, you feed it into those machine learning models, and then you store the results in vector format into Chroma.

Jeff Huber [8:51] That's right. So the workflow is: take your documents. This is the most classic workflow, if you will. Take your documents. It could be, again, your knowledge bases, your documentation. You saw a demo of WeWork's S-1, whatever. Break it into small pieces. Pass it to the embedding model. The embedding model turns each small piece of text—you can also do images, but let's use text for this analogy—into a number. And then all of those numbers and all of that text gets loaded into this database and becomes quickly searchable.

Jeff Huber [9:12] Which is the fundamental thing that you need to be able to do, is not just brute force through the whole thing, which can be quite slow at scale, even smaller scale. But what you want to be able to do is get response time in tens of milliseconds.

Matt Turck [9:31] Okay. So what does Chroma do today? So obviously, the core mission, as we just discussed, is to store and process embeddings, but there's a list of features in terms of what you do now, maybe what you will do soon. What are some of those key aspects?

Jeff Huber [9:52] Yeah, I think our focus is around developer experience. So all of the work that we've done thus far is trying to meet developers as early as possible in their lifecycle and make the technology simple, understandable. You go to the website, even the design of the website is meant to be calm, and, I promise, you can learn this. It's not that hard. It's really not hard. So the kinds of features that we're excited about, so the things that we're working on broadly, we're working on a distributed version of Chroma.

Jeff Huber [10:28] So in the same way that Elastic, for those of you that are familiar with text search, picked up Lucene, made it developer-friendly, and made it distributed, Chroma picks up some of these ANN algorithms. We can go into more detail about what that means. We've made it developer-friendly, and now we're making it distributed, again, all in open source. That's on the scale vector. And then the other vector is making it usable and understandable. It's basically 2D data. You can put it on a screen.

Jeff Huber [10:56] You can immediately understand what's going on. It's extremely intuitive. It's 2D data. It's a table. You've all used Excel. This data is not 2D. It might be hundreds or thousands of dimensions. We think that, again, there's a large gap to make this technology really usable and interpretable to humans, to developers who are building these applications. So there's a large swath of features and functionality which we're excited to bring to developers. Just to give you one example of this, and again, a lot of this work is still fully in open source.

Jeff Huber [11:26] One of the things that developers who adopt this technology or related technologies face, these questions they face are questions like: How do I chunk up my document? How large or small should these pieces be? Which embedding model should I use for my application? How many nearest neighbors should I be picking? Because that's kind of a feature of this task. You have to pick: Do I want three nearest neighbors or 10? And then, are these nearest neighbors relevant or not?

Jeff Huber [11:52] For all of these questions, this is the classic workflow every developer faces. The best answer today is, just try it out. We have no idea. And we think that there's a large opportunity as well to bring applied research to bear on that problem and give developers better tools to be able to answer these questions and build more reliable applications more quickly. Very good.

Matt Turck [12:01] You are very proudly open source. Do you want to talk about this?

Jeff Huber [12:02] Why?

Matt Turck [12:03] How do you think about it?

Jeff Huber [12:22] There's all kinds of strategic reasons, et cetera, et cetera, you can get into. Maybe in the context of a VC pitch that would be relevant. I think the honest truth is because it's awesome and it's really fun, and I want to do things in my life with people that are awesome and fun. And so that's why. I think on the comp side, there's a long history of open source databases that have also been able to monetize well and monetize in a way that the community loves.

Jeff Huber [12:47] You can look at MongoDB, Elastic, many examples of this. So I think the monetization story is really aligned with the community, which is maybe unique, right? Other open source products, it gets a bit more contentious.

Matt Turck [12:52] For the open source geeks out there, what license do you have?

Jeff Huber [12:53] Apache 2.

Matt Turck [13:01] Apache 2. Okay. And you alluded to a cloud product. Is that coming up, or do you have it already?

Jeff Huber [13:04] Yeah, it's not live yet. We're working hard on it.

Matt Turck [13:05] Okay.

Jeff Huber [13:30] Yeah, the reason we're doing it is not to really make money. The reason we're doing it is because it's what the community wants. The meme internally has been, "When hosted?" And unsurprisingly, developers love building stuff. Especially application developers, web developers, mobile developers love to build. They don't really want to manage their own infrastructure. And so while they love the experience of using Chroma on their own computer, we don't have a good story yet for how they go to the cloud or to production.

Jeff Huber [13:52] And so part of the idea with Chroma is it'll be the same API everywhere, from running it in a Jupyter Notebook, running it on your local computer, or running it in the cloud. And so that's—

Matt Turck [13:58] Okay. So you're going to have a self-hosted version. So you're going to have open source, cloud, and sort of self-deployable, self-hosted?

Jeff Huber [14:26] Yeah, of course. I mean, Chroma will always—we are committed to building the ubiquitous open source standard. And so a truly non-kneecapped open source database has to exist. I think that that's A, the ethically right thing to do, and then B, it's honestly pretty strategic as well. You see projects that try to have it both ways, try to, like, "Oh, well, you can host it yourself, but it doesn't have these three critical features. Good luck." I don't know. I just think it's dumb.

Jeff Huber [14:28] Okay.

Matt Turck [15:00] The vector database space has been super fun to watch recently because it's a term that hardly anybody knew, what feels like a year ago. And suddenly there seems to be a number of players that are claiming to be great vector databases, and maybe that's all correct. How should we think about all those companies? Are you all largely going in the same direction, and it's a race to whoever builds the most features the fastest, or are they different approaches?

Jeff Huber [15:26] Yeah, I mean, I think I respect all of our competitors, all of our competition. I think everybody means well. I think everybody wants to do the right thing, but I'll speak to what we care about, and what we care about is developer productivity and happiness. And I think that's hopefully evident in what we've made thus far, and hopefully will be true about what we will do in the near future and long into the future.

Matt Turck [15:51] And how should one look at this if you are a developer and you're comparing solutions? Is that a question? So there's developer experience, so how easy it is, documentation, all the things. Is there a concept of performance? And if so, how do you measure it for a vector database? What are the various axes for one to make a selection?

Jeff Huber [16:14] Sure. Yeah. I think there's, again, nobody knows the future, certainly not me. There's an interesting open question, which is: What will the primary workloads look like? Try to avoid using overly database jargon. What will the workloads look like? Will they be more transactional, meaning it's sort of a classic database, you're constantly querying it and updating it, deleting stuff? Will it be more analytical, where you want to slurp in the whole database and then do some analytics process over the whole thing?

Jeff Huber [16:37] Or will it be both? Will actually both be important? And we fall in the latter camp. So the acronym here, again, for the database geeks would be HTAP: Hybrid Transactional Analytical Processing. And we think that both—certainly transactional—has to be the case because it is an online database. It's going to sit in the loop of applications. Again, you've already seen demos of this happening tonight, but also to make this technology useful for developers, you do want to slurp in the whole database and analyze it in a bunch of different ways.

Jeff Huber [17:07] And so that's an analytical process. And so again, to not go too deep on database jargon here, but we think that both are really important. And I think that again, without getting into some details, that's unique, that we have a unique perspective on that.

Matt Turck [17:19] Okay, great. And talking about go-to-market a little bit, what are some of the early learnings and lessons in terms of the range of use cases that people use Chroma for?

Jeff Huber [17:48] Yeah, I think the most common use case is what you would think. It's the chat-your-data use case, which again, you've seen a bunch of demos already tonight. And on one level, I think it's fair to trivialize that and say, this is the newspapers-on-the-internet moment. Oh, the internet came. What are we going to do? Let's put newspapers on the internet. It's just sort of like pattern matching. We already know how to do this. Now let's do it in this way.

Jeff Huber [18:21] And I think that the secondary observation is that's not wrong. All newspapers are on the internet now and arguably are much more successful because of the internet. So I think that use case, even just the chat-your-data use case, truly will go to the ends of the earth. I think the more new use cases you're starting to see emerge is the ability to give LLM software agents their own memory, especially in the context of multi-agent interactive things. So you've seen things like BabyAGI or AutoGPT. There was recently a project called Voyager that used Chroma on the backend for the vector storage and retrieval.

Jeff Huber [19:03] That I think is really interesting and is native to AI. You couldn't do that before AI. So it's truly native to AI, and that's really exciting. Now, of course, for those of you that have actually played with this technology, I think it's questionable whether the current state-of-the-art language models, embedding models, et cetera, will give you the reliability you want from agents working together. But it almost certainly seems like the future.

Matt Turck [19:25] And on that point of reliability, a point that is very important in all of this discussion that we haven't really quite made yet is that vector databases are a way for companies to limit the hallucination problem. Do you want to just double-click on that?

Jeff Huber [19:49] Yeah, there's maybe some nuance here which is interesting. Hallucinations are when you ask a language model to do something, it kind of ignores you and it just makes something up. There's a famous case recently where a lawyer was caught using it in a brief to a judge, and the judge rightfully shamed him. The solution to hallucinations is to give the language model the information you want it to have, again, through this vector search, vector retrieval process that we've been talking about tonight.

Jeff Huber [20:22] It's a grounding task. Literally, the prompt construction that you typically see here is, "Hi, language model, please do not make anything up. If you don't know, say I don't know. Only consider the following documents: one, two, three. Now answer the following question: question." That's the prompt that's being built on the backend, and how those documents from vector search are being inputted, and then how the prompt from the user is getting slotted in. That's the state of the art today.

Jeff Huber [20:34] There's lots of stuff around steering language models at the embedding layer itself and not using text, but this is not yet exposed to at least closed-source models.

Matt Turck [21:08] While we're on the topic of how enterprises and companies use vector databases, where does that fit? That seems to be an emerging stack around foundation models and the LangChains of the world. Who does what? And if you're Moody's today, which is kindly hosting us, and you want to build your own generative AI stack, so clearly you use Chroma, but what else do you need to have?

Jeff Huber [21:32] Yeah, again, I think emerging is the right label. And so we will see. Reasoning by analogy, really building with language models is a new way of building software. It's a way of building software that would've been impossible to program by hand. Just by simply typing into the black box of a language model, "Hey, language model, here's a paragraph of text. Tell me whether this is happy or sad." That's something that's so easy to do now with the new primitives that we have available that even just a year ago would have been extremely difficult.

Jeff Huber [22:11] Extremely difficult. Would take an entire research organization years potentially to build. I mean, maybe not quite that complex for a simple semantic classification task, but still much more difficult. So for this new way of building software, again, it won't necessarily eat all existing ways of building software, but it's certainly a new way of building software. A stack will emerge. And I think again, reasoning by history, by analogy, what do you have? You have databases. It's where you store stuff.

Jeff Huber [22:32] You have application code. It's where you encode your business logic about guiding the task, what you want it to do. You have your inputs and your outputs. You've got some UI. I'm not sure how much that's going to change with AI. And then I guess the LLM, in my mind, is sort of like a super business logic black-box plugin. So I think that basic model will probably still hold, where you have things that look like databases pretty much, things that look like application code pretty much.

Jeff Huber [23:10] Again, the language model is sort of a new thing in some ways, and then it'll be different than it has been before as well. Reasoning by analogy, obviously you have to be careful with how far you take it. We think the database will be much thicker than it's been before. Postgres is not doing analytics over your Postgres data to help you understand your Postgres data. I've already talked about tonight the idea that I think that's really valuable and important. I think obviously, too, language models will be in the loop of a lot of these things.

Jeff Huber [23:45] And so whilst we want to try to separate them cleanly and have these clean lines and abstractions, everything truthfully will be more shades of gray in between all of these components. And there'll be language models running inside the application code, obviously, language models running inside the database as well, or large models more broadly. Multimodal models will run everywhere as well. So it probably won't be super clean, but I think it'll be more or less like what we've seen before.

Matt Turck [24:12] Great. One last question from me, because then I want to open up to folks. Drew over there has a microphone, which hopefully works. So zooming way back out, where do you think we are in a year or two from now in generative AI? Is that just deployed everywhere? What's your vision of the kind of short-term future?

Jeff Huber [24:41] I think the long-term future, so maybe three to five years from now. Two years is this dead zone of prediction, right? You can predict the next week and you can predict the far future, but don't try to predict two or three years out because you're always going to be wrong. So the farther future is intelligence too cheap to meter. Intelligence built into every product, every service that we use. Intelligence too cheap to meter. Hopefully leading to a lot more flourishing for all of humanity alongside clean energy too cheap to meter.

Jeff Huber [25:23] So that's the big vision. Where will we be three years from now? I certainly think probably most enterprises, organizations, companies on Earth will have brought language models into the company, probably in pretty meaningful ways. At minimum, the customer service department, the sales department, ops, backend, legal, and hopefully also the core product experience that they deliver to users, to their customers. So that seems like a safe bet. So I'm willing to make that bet. Great.

Matt Turck [25:25] All right, very good. That's a wonderful place to leave it. Thank you very much.

Jeff Huber [25:39] Thanks, Matt. Thanks for listening to The MAD Podcast. If you liked this episode, be sure to leave us a review. com/events/data-driven.