Vectara: LLM Powered Search with CEO Amr Awadallah

The MAD Podcast with Matt Turck · with Amr Awadallah, CEO, Vectara

Amr Awadallah is the CEO at Vectara. We cover why grounded generation retrieves facts before generating answers to minimize hallucinations, why fine-tuning private data is slower and 100 times more expensive, and why action engines require zero hallucination before software can safely execute requests.

Watch on YouTube

Chapters

  1. 0:00 — Full episode

Transcript

Full episode

Amr Awadallah [0:40] Hello and welcome back to The MAD Podcast, a series of conversations with leaders from across the data, AI, and machine learning landscape by FirstMark Capital. Today we have a great episode about LLM-powered search and AI hallucinations with Amr Awadallah, founder and CEO of Vectara, and Matt Turck, managing director at FirstMark. This episode was recorded live at a recent Data Driven NYC event. Good to see you again.

Matt Turck [1:15] Thanks for doing this. We were just reminiscing right before this that you spoke at this event in February 2020, and that was the last event before the pandemic. Yes, and we had to cancel the March event, but I'm very excited to have you back. At the time, you're the original founder and CTO of Cloudera, so a very major company and very major success in the data infrastructure world. That talk was with your Google VP hat on.

Amr Awadallah [1:16] Yes.

Matt Turck [1:49] And now you're back at it. You've been back at it with Vectara. You're the CEO. This is an LLM-powered search company. You'll explain what that means. And you came out of stealth, I guess, in October, announced a $20 million seed round as well. So I'd love to start from the top and maybe explain what the whole LLM-powered search means versus keyword search and existing solutions.

Amr Awadallah [2:10] Yeah, so it was very hard to explain what we do, actually, when we launched in October. It was really hard to explain what we do. And then ChatGPT came out, and then it became very simple. So what we do is ChatGPT for your own data, right? So imagine yourself: you have an intern working for you, and this intern—

Matt Turck [2:11] Writing my tweets.

Amr Awadallah [2:12] Say again?

Matt Turck [2:13] Writing my tweets.

Amr Awadallah [2:38] Yes, exactly. And that intern is multilingual. They speak tens of languages. And you give them all of the PowerPoint decks that you ever receive. You give them all of the transcripts of the calls that you convert into text using the amazing technologies that Dylan just talked about a second ago. You give them the executive summary reports, the investment memo reports you write yourself when I approve or decline, and they get that in multiple languages. You might have another department in France doing it in French, in China doing it in Chinese, and in Japan doing it in Japanese.

Amr Awadallah [3:08] And then you have a conversation with that intern, and you tell the intern, "Last time when I had somebody pitching me for blockchain from Europe, why did I decline? Why did I pass?" And the intern would tell you, "Oh, I saw that report that was written in French, but now I'm going to give you a response in English because you, Matt, you only speak English. The reason why I declined is one, two, and three." And then you tell the intern, "Oh, tell me more about the second reason."

Amr Awadallah [3:31] What does that mean? And they explain to you the second reason and what it meant. And then you tell them, "If this was an entrepreneur from the U.S. instead pitching me on blockchain, with everything else equal, would I have still said yes?" And they would tell you the probability you would have still said yes is 75%. So essentially, you're now having a conversation with your data, where the data is an active participant with you in the decision-making that you are trying to achieve in your organization.

Amr Awadallah [4:06] And that's really what we're after here. To do that right, you have to solve a very key problem with large language models that some of you might have heard about. That's called the hallucination problem. And the hallucination problem is where sometimes, because of the probabilistic, stochastic nature of these models, they make up shit. They literally just make up shit. And that can't be. You can't say, "What were the three reasons I declined?" and it gives you one, two, and then the third one is something it made up.

Amr Awadallah [4:20] That's not gonna work. That's not gonna fly. I mean, you do have a human in the loop, so you can still qualify it, but we want it to be minimized to zero. And that's the problem that we solved.

Matt Turck [4:47] So if I was to deploy Vectara in my enterprise, I would basically connect to lots of different sources of data. And in the traditional world of search, you would go and index the content. Is this completely different, or is there some overlap between the way traditional search works and the way neural search works?

Amr Awadallah [5:12] Very, very, very good question. So the previous wave, or the classical way for how we did search, all of us, and the way we have been trained on it with Google and many other search systems that we have been using, is what's called keyword search, right? That's where we're trying to match the keyword you're looking for with what's in the content. So you search for weather balloon, and we try to exactly match weather balloon. We don't try to match the meaning of weather balloon.

Amr Awadallah [5:38] The meaning of weather balloon should also match UFO. Nowadays it does match UFO; it wasn't before, but now it does. But an intern that has a smart brain that understands the meaning would do that matching. So these newer techniques for doing information retrieval, they are based on, again, a neural network model that is almost as smart as us in terms of mapping from language space to a meaning space. And the beauty of these models is they can do it across languages.

Amr Awadallah [6:05] So whether that is weather balloon in English or Chinese or Japanese or Korean, all of them in the human language space, the neural network maps it to the same meaning in the meaning space. And that now allows us to always fetch the proper meaning for what you're after. What we provide at Vectara is an API that allows you to do that. So we sell to developers. We have an API that allows you to upload any type of content. We have only two API calls.

Amr Awadallah [6:31] Actually, there's a bunch more, but you can think of it as two API calls. There is one API call to upload the content that you have. That content could be text documents, could be JSON, XML, PowerPoint, Excel, you name it. And then another API call where you're issuing the conversational search aspect of it, and then it just gives you back the response that gives you back an answer. It doesn't give you a list of results anymore. Again, keyword search trained us with lists of results.

Amr Awadallah [6:58] Result number one, result number two, and you have to go read them to figure out what the answer is. If you have something that comprehends and understands, it just gives you the answer. So we're moving away from legacy, which I'm calling search engines—that's legacy—to what we have today, which is answer engines that just give you the answer itself. So you don't have to go and parse all of that content. It's very similar to Bing Chat. Can you raise your hand if you use Bing Chat?

Amr Awadallah [7:06] It's amazing that actually there's hands coming up. Like, you would ask people about Bing...

Matt Turck [7:08] The biggest surprise of 2023.

Amr Awadallah [7:25] So yes, Bing Chat, when you ask the question about the web content—so this is not for your own data now, this is for web content—it gives you the answer and it gives you the citation. Like, in the answer, this first sentence came from this citation, this second sentence came from this citation, this third sentence came from this citation. So we do the exact same thing. So when you ask a question, we will give you the answer, but we'll tell you this came from document number five over there.

Amr Awadallah [7:37] This came from the PowerPoint slide over here. This came from the email over there. And you know exactly where that answer came from.

Matt Turck [7:50] Great. So back to the hallucination problem. You have built a novel approach that I believe you call grounded generation. I would love to hear what that does.

Amr Awadallah [8:13] Yes. So this approach, I don't want to take credit for it. There's many researchers that are advocating for this approach. It's called grounded generation. So you want to generate, but the generation is grounded in the facts. So you're not telling the neural network, go give up an answer from whatever you have learned in the past. No, you're telling it, constrain your response to the facts that we're giving to you. Another way you will hear it referred to is retrieval-augmented generation.

Amr Awadallah [8:39] So you're retrieving first. You have a very good neural network that specializes in retrieving the facts, and then the generation is a function of these facts. So before I talk about the technical stuff, I want to step up a level here because I'm assuming there's a wide audience here that's not engineers, where I want to go back to humans first, right? So imagine you, Matt, and I gave you 10 books, okay? And I tell you I'm going to be asking you questions about these 10 books, right?

Amr Awadallah [9:09] There's two approaches you can choose to do that. The first approach is you will actually read all of them books, comprehend them, understand them, reconfigure your neural network up here to encode the contents. And unless you have a photographic memory, you're not gonna remember everything in these 10 books, right? Which means you will hallucinate. So when I ask you a question about these books, you're gonna try your best to recall what was in there and answer my question. But you're not gonna be able to recall it in the exact right way, and that's where hallucination creeps in.

Amr Awadallah [9:38] So grounded generation, or retrieval-augmented generation, is a different approach. It's you, and you have an intern now beside you. When I ask the question, the question goes to the intern first, and the intern very quickly highlights with a highlighter, imagine a yellow highlighter, highlights the sentences that are most relevant to that question I'm asking. It's not the entire book, it's the sentences that are most relevant to the question. And then they give it to you, Matt. Now the genius that knows the language and knows maybe financial industry text or knows legal industry text, you look at these facts and you say, aha, this is the perfect answer.

Amr Awadallah [10:16] So that's what we call retrieval-augmented generation. To be more technical in detail, there are three models that get applied in doing this. The first model is the encoder model that knows how to go from language space to meaning space. My co-founder at Vectara, his name is Amin Ahmad. And Amin worked at Google Research, and he was one of the key inventors of a technique that's called the Multilingual Universal Sentence Encoder. This is a neural network that knows how to go from all human languages into a unified meaning space.

Amr Awadallah [10:40] It's almost like a lingua franca that the neural network came up with. And the amazing thing, by the way, about this neural network: even if it has not seen a language before, it can still do a good job at it. So we just trained a new model, a fresher version that we're working on right now, on about 30 languages, and Vietnamese was not one of them. And it still figured out how to go from Vietnamese to—it's amazing, these emergent behaviors of these neural networks.

Amr Awadallah [11:08] So that's the first model, right? So the first model goes from human language space to meaning space, and that's how you can do the retrieval to find the facts. But once you get the facts, you want to order them in the right order. You want to have the most relevant fact to the question be number one, the second most relevant be number two. For that, we use something called a cross-attentional re-ranker. It's a BERT-like model that is able to re-rank these facts as a function of the question.

Amr Awadallah [11:38] Once you finish re-ranking the facts, then you give it to a summarizer model, which is just similar to GPT, and you tell the summarizer model, please read these facts. You know the English language very well or the French language very well. You know the domain very well. Summarize these facts now into a response, but constrain your answer to only be a function of these facts. Don't go hallucinate stuff from your brain that you have remembered from before. Constrain your answer to what the facts are telling you.

Amr Awadallah [11:55] And that's how you minimize the hallucination problem. You minimize, you don't eliminate completely. Every now and then it still can insert something. And we're working very hard to make it 0% in terms of hallucination. Hopefully that communicates the meaning.

Matt Turck [11:57] Yeah, that's super interesting.

Amr Awadallah [11:59] And hence grounded generation. It's grounded in the facts.

Matt Turck [12:11] So is that beyond Vectara? Is that the future of how LLMs get deployed in the world where the quality of answer matters?

Amr Awadallah [12:32] I think there are two ways to have GPT for your own data. The first way is to do what's called model fine-tuning. That is, here, Matt, go read all of these 10 books and relearn all of these concepts in your neural network. I think that way is susceptible to hallucination. That way is very expensive because retraining a neural network actually is very costly. That way is very slow. A new fact coming in, a new book coming in, you have to retrain.

Amr Awadallah [12:56] It's not going to be available in the outputs until weeks later. The grounded generation approach is real-time. A new fact comes in, it's showing up in the answers right away. There is no hallucination. I mean, you minimize it significantly, and the cost is 100 times cheaper. I am biased because that's what I'm selling, but yes, I think that is the new way. The new way is going to be the combination of excellent retrieval models that know how to fetch the facts and then domain-specific models, whether that be legal or finance or health, that know how to take these facts and convert them into a response for the question the user is asking.

Matt Turck [13:21] How experimental is this versus we already know it works? You mentioned that people have been thinking about this for a bit.

Amr Awadallah [13:37] Yeah, no, this is production. I mean, we had this in production since October. Bing Chat, that's exactly how Bing Chat works. So Bing Chat, when you use Bing Chat, and I invite you to go try it today, it searches the web first. If I ask Bing Chat, who is Matt Turck? Matt Turck, right? Is that the right way to say your last name? If you ask GPT and it doesn't know you, it will make shit up about you, right?

Amr Awadallah [13:44] It'll say Matt Turck is an amazing violinist, which is, I don't think you can—

Matt Turck [13:45] Well, that too, but—

Amr Awadallah [14:06] Right, but what Bing Chat would do, no, it would go search first. It will find the facts on the web about Matt Turck, and then because it knows the English language and how to convert these facts into a response, it will display the answer. This is already how it's working today. So this is, no, this is not experimental. This is how it's being done. Now, how do we minimize hallucination further in the future? We want to minimize it even more, is to do what's called fact-checking.

Amr Awadallah [14:34] And we don't have that yet. But most media outlets like The Wall Street Journal and The New York Times, the reporters, when they finish writing an article, they still have to go to a fact-checker that checks the facts in the article they've written because reporters can hallucinate. They're susceptible to that. And the fact-checkers would look at some of the key facts and make sure they were grounded in facts and, in fact, they are true, and then the article gets approved. We haven't done that yet, but I think that's the next step of how you're going to have models that check the accuracy of other models, and that will help us bring down hallucination to zero.

Amr Awadallah [15:00] I'm sorry to interject here. One amazing thing: if we can bring down hallucination to zero, to be 0%, then now we can go from search engines to answer engines, which is where we are today, but the human is still in the loop. And then the future is going to be action engines, where the action can be taken. So, for example, if I'm trying to take a vacation in an HR system—actually, show of hands, how many of you use Workday at your job as one of your HR systems?

Amr Awadallah [15:29] About 30% of the room raised their hand. Workday is so hard to use. If you want to take a vacation, you have to fill in many forms and click on many buttons. Every time, you have to read the documentation to figure out and remember how you should do it. Imagine now you ask the system, "I would like to take a vacation," and it just responds back, "When are you leaving?" And you say, "I'm leaving on this day." "When are you coming back?"

Amr Awadallah [15:45] And you say, "I'm coming back this day." "Do you want to use your paid time or unpaid time?" And you give it the answer, and it says, "Okay, these are the steps." And then a button shows up and says, "Would you like me to do that for you?" And you just click on it, and it gets done. But we can only do that when you bring down hallucination to zero, because what if one of the steps is, "And I'm now resigning," and it just resigns on your behalf?

Amr Awadallah [16:08] Then that wouldn't work. So I think as soon as we bring down hallucination to zero, this is gonna be an amazing world. The way we interact with software is gonna be verbal, whether that be written or spoken, and our lives will be much easier than they have ever been.

Matt Turck [16:21] Great. So back to Vectara as a business: how do you sell? Who do you sell to? What are the use cases? How does the business side of the story work?

Amr Awadallah [16:43] Yeah, excellent question. So I don't know if Will is still here, because Will was talking about open source. I have been through open source, and I have the battle scars from open source and how brutal Amazon is at competing with open source. Open core does not work against Amazon because open core, you're building the core open, and then manageability, security, reliability—that's your proprietary... Amazon is really good at that shit. Like, they know how to do security.

Amr Awadallah [17:00] So it becomes very hard to compete with them. So I can give him some tips and lessons there. If you look at the most successful software company that IPO'd in the last 10 years, who would come to your mind as that? The company that IPO'd that was most successful valuation-wise.

Matt Turck [17:01] Snowflake.

Amr Awadallah [17:22] Snowflake. 100% closed, no open-source anything in Snowflake. It's a 100% closed platform, but it's easy to use. It solves the customer problem, right? It's very easy to start with a small amount of data. You only pay them $3K per year, and then you land with them, you get hooked on how amazing they are, and you grow with them. And that's exactly the business model that I'm adopting at Vectara, right? I will still do open-source stuff because I'm open source at heart.

Amr Awadallah [17:48] And I love open source as well, just like Will does. So I will do open-source contributions, but my core platform is only available as a proprietary API, just like Snowflake. The business model is very simple. We charge in the same way that Snowflake charges, or Google BigQuery, or Amazon Redshift. It's a function of how many API calls you're making against the platform and the amount of data that you have in your platform. And the business model is about: we land you, even if it's for a very small amount of data.

Amr Awadallah [18:13] We show you magic. You get blown away by how easy this is compared to doing anything else, and just start adding more and more data and more and more use cases. So what are the use cases? I don't know what the use cases are going to be in the future. In the same way, when the iPhone moment came out and it enabled us to use our fingers to do stuff, we didn't think of Snapchat. And then Snapchat and swipe right and swipe left, Tinder—that came later, right?

Amr Awadallah [18:39] I think there will be many new applications that will come out of this. But I know what I'm selling today. I'm selling two use cases. The first one is customer support. Customer support is amazing for this because when you are complaining, "My washing machine is acting up, and this noise is coming out, and this light is coming on the front. What should I do?" That is semantic. That's not keyword lookups. Or you say, "Oh, I'm using Snowflake."

Amr Awadallah [18:45] "The queries are too slow, and I have this index. What should I do?"

Matt Turck [18:45] Again, semantic.

Amr Awadallah [19:08] It needs something like this to give you the answer. I don't want to go and read manuals and manuals and manuals and knowledge-base articles to figure out how to do it. So the number one use case by far right now is customer support. The number two use case is what we refer to as knowledge discovery. So knowledge discovery is the financial analyst trying to make a decision based on historical decisions they made in the past. So hedge funds and finance folks.

Amr Awadallah [19:30] The lawyers doing discovery in a lawsuit or doing a contract evaluation and looking at previous contracts and how they match this condition or not. The media outlets that are trying to track whether balloons, whether they're being mentioned in the news, or a given CEO saying something in their earnings call in this way or that way. That's another very, very common use case. The pharma companies looking at the research coming from the trials that they're doing for the drugs and looking at adverse conditions.

Amr Awadallah [20:01] But adverse conditions can be written and described in different ways across different languages. We find that right away in the right way. So these are the two primary use cases we're focused on. There is another good one, which you alluded to, which is called enterprise search. We are actually avoiding that one right now. We might come to it in the future. Intranet, from SharePoint, and then you assemble them all in a nice UI across your organization. That's number 11 on the CIO list of things to spend on.

Amr Awadallah [20:15] I'm not going to go after that during an economic downturn. Maybe later on we'll approach that. I think it's an important problem to be solved, but it's not the primary one we're focused on right now.

Matt Turck [20:29] But again, it's all API-based so that they build—they being customers—build it into their internal or external applications.

Amr Awadallah [20:58] Excellent question. It is. But we also have reference open-source implementations for how to do an answer service just like Bing Chat, how to do a conversational bot that converses with the data, how to do semantic matching for triggering alerts when certain semantic meanings happen. So all of these open-source implementations, we give them to our customers for free so that they can accelerate the deployment of our model. Anyway, go ahead.

Matt Turck [21:09] Sorry. Let's talk about the building blocks of the Vectara platform. You alluded to some of this, but I'd love to hear the full story.

Amr Awadallah [21:24] Yeah, excellent question. So again, our job is very similar to what Dylan mentioned, the speaker just before me: he wants to make this simple for the average developer. I always like to break down developers in the world, and I learned this the hard way at Cloudera, my previous company. I came out of Yahoo, my co-founders for Cloudera came out of Facebook and Google, and we thought that the developers in the world are all like the developers we saw at Yahoo, Facebook, and Google, and they're not.

Amr Awadallah [21:57] Right? So, the developers in Silicon Valley, and many of those also you see here in New York as well, they are what we refer to as descriptive developers. Or, in layman's terms, I like to call them the Home Depot developers. So, what does that mean? If they say, give me the planks of wood, give me the nails, give me the hammers, give me the drills, I want to build a desk, right? But I'm going to figure it out. Don't tell me how to build a desk.

Amr Awadallah [22:17] Tell me what the things can do, and I'm going to put them together and make the desk. So we call those descriptive developers. They only want you to describe to them what you do. Don't tell them what to do, right? Which is great. There's a good market for that, and there's companies that sell to that, and they will make some good money from it. The rest of the world, the entire world, Indonesia, the U.S., Ohio, like different states, they are prescriptive developers.

Amr Awadallah [22:48] They prefer the IKEA model. The IKEA model is, here is the recipe for how to put the desk together. Put these pieces, put these screws over here, you're gonna have an amazing desk on the other end. So in the case of voice transcription, here is an amazing API. We're gonna take care of the model, the right model for you. You plug in the API, one side you put the voice, other side you get amazing text and amazing sentiment. In our case at Vectara, we give you a very simple API.

Amr Awadallah [23:09] On one end you plug in the data, on the other end you get amazing outputs in terms of conversational search. The components in the middle, you don't have to worry about at all. But that said, if I want to uncover what's in the middle, as I mentioned, we have one model that's good at doing the encoding from language space to meaning space. We have a vector database that does the vector matching to find the right vectors that match the meaning that you're after.

Amr Awadallah [23:31] We have a cross-attention re-ranking model that re-ranks these facts so the most relevant fact is first, and then we have a summarizer model that provides the final answer. And the summarizer model can be different ones. It can be one for legal, it can be one for finance, and we pick the right one for you. Just like Twilio picks the right network for you, we do that automatically for you so that you, as a developer, now can be a hero overnight without having to learn all of this stuff.

Amr Awadallah [23:48] We take care of making this stuff work perfectly for you. So that's our approach, and these are the components in our platform.

Matt Turck [24:17] So you just mentioned vector databases, and a lot of us come to these events to just learn about the space and try to figure out what company and product fits where. So in a world where you have model producers like OpenAI and then vector database producers like Pinecone, which, by the way, is speaking at the next event, LangChain, all those things, where does Vectara fit?

Amr Awadallah [24:39] Yeah, so a vector database like Pinecone, that is one of the building blocks. That's like the drill at Home Depot, right? So if you are one of the developers that wants to build your embeddings yourself and put the LEGO blocks together yourself, then you would use Pinecone. Then you would use LangChain to build a pipeline. And you want to understand the different models in the pipeline and how you're going to put them together. And please, go use that. And actually, for prototyping, it's great to use that.

Amr Awadallah [25:07] Now get that into production at scale, and that's when you're going to start running into hiccups in terms of scalability cost-wise, scalability performance-wise on the ingest side, and then on the query runtime side, how we can get back the answer in a few milliseconds as opposed to hundreds of milliseconds. And that's really where we shine. So if we start talking to a developer and we start getting the developer asking us, why aren't you using this model or that model that we saw on Hugging Face? We tell them, okay, you're not for us.

Amr Awadallah [25:26] You go work in Hugging Face. I want the other developer that wants to focus on solving a problem for their company and making their company more efficient working with these things. I don't want to insult anybody who likes the other approach. That's great, and it's an amazing approach. It's not who we are after.

Matt Turck [25:40] Again, like zooming back out, what do you make of the whole pause idea in research, stopping all AI research for six months? Does that make any sense to you?

Amr Awadallah [26:04] So first, show of hands, how many of you heard about this letter asking for a pause in research for six months? So 50% of the room raised their hand. So this letter came out a few weeks ago and a number of people signed it, including Elon Musk. I think it's mainly Elon Musk wanting to catch up. So he's asking OpenAI and everybody else, just pause until I buy all of my GPUs and catch up with you. I'm half joking there.

Amr Awadallah [26:30] I actually partially agree with the spirit of the letter. I don't think we should pause anything. By definition, whenever you have economic efficiencies like the Industrial Revolution or the mobile revolution, or now this amazing GenAI revolution, we're not going to stop. We're going to make it happen. We're going to make it work. We're going to figure it out. We know that's how capitalism works. However, the caution of being careful how these technologies can be abused is definitely something I'm very, very concerned about.

Amr Awadallah [26:56] So today there are cloning technologies that can clone our voice very easily. I can clone your voice right now, call up your wife and tell your wife I'm really stranded in this corner by Broadway and whatever, and she will show up thinking it's you. And there's nothing that prevents me from doing that today. What these folks have a responsibility to do is to make sure there is a very clear voice signature where you, Matt, are saying, I approve, with your own voice.

Amr Awadallah [27:31] I approve that this other voice, which is my voice, and the match is 100%, can be used in this way. So that's one example. There's many other abuses taking place, much worse, by the way. Like, I don't want to discuss them here. I 100% support that we need to have regulations that prevent us from shedding the responsibility that we have to make sure that these technologies are not used for harm. Thank you. So that I am for.

Amr Awadallah [27:34] Pausing research, absolutely not for.

Matt Turck [27:48] All right, well, that's a great topic to talk about over pizza and wine and beer. So please join me in thanking Amr and all our speakers tonight.

Amr Awadallah [28:02] Thanks for listening to The MAD Podcast. If you liked this episode, be sure to leave us a review. firstmark.com/events/data-driven.