Glean’s Breakthrough: CEO Arvind Jain on Scaling AI Agents & Search
The MAD Podcast with Matt Turck · with Arvind Jain, CEO, Glean
Arvind Jain is the CEO at Glean. We cover why enterprises start with closed models before distilling open-source models at scale, why agents still require human supervision despite reducing contract redlining from two weeks to one minute, and how Glean ranks enterprise knowledge using permissions, recency, and human behavior.
Chapters
- 2:01 — The AI model explosion: open vs. closed in the enterprise
- 6:19 — Why enterprises choose open source AI (and when)
- 10:33 — The agent era: what are AI agents and why now?
- 12:41 — Automating business processes: real-world agent use cases
- 16:46 — Are we there yet? The reality of AI agents in 2025
- 19:24 — Glean’s origin story: reinventing enterprise search
- 26:38 — Glean agents: from apps to agentic platforms
- 31:22 — Horizontal vs. vertical: Glean’s strategic platform choice
- 34:14 — How Glean’s enterprise search works
- 39:34 — Staying LLM-agnostic: integrating new AI models
- 42:11 — The architecture of Glean agents: tool use and beyond
- 43:50 — Data flywheels and personalization in Glean
- 47:06 — Moats, competition, and the future of work with AI agents
Transcript
The AI model explosion: open vs. closed in the enterprise
Matt Turck [1:17] Arvind, welcome to The MAD Podcast.
Arvind Jain [1:18] Thank you for having me.
Matt Turck [1:49] This feels like a very timely moment to have the conversation about Glean. It seems that you guys are firing on all cylinders. You had a well-publicized, remarkable year in terms of ARR acceleration. You made, just a few weeks ago, what feels like a really important release around agentic reasoning. And of course, there are rumors, which I will not ask you to confirm nor deny, that there may be another big round afoot with a rumored $7 billion valuation. So thanks for making the time in the middle of all this to chat with us.
Matt Turck [2:17] We're going to unpack all things at Glean: enterprise search, RAG, agents. But before we do that, I thought it'd be a fun way to start to sort of riff a little bit on what's going on in the industry. So it's been yet another wild week in AI: Gemini 2.5 Flash. So I'm curious, as an industry participant and observer, what do you make of all of this? And where do you think we're heading in terms of what the AI model world is going to look like in a few years?
Matt Turck [2:33] Are we going to have a bunch of models? Are we going to have a small number of very powerful models and model companies?
Arvind Jain [3:13] I think all indications are that there's going to be a very large number of models that are going to be available to us in the industry. It's going to be both closed models from companies like OpenAI and Anthropic and Google, but also there's going to be a lot of open-source models that are like, for example, Llama, and then a lot of derivatives as well that are going to be available to us. So that is already happening. It's just going to be more and more of that in the future.
Arvind Jain [3:53] What it means for us is we will be in this world where it's going to be hard to actually always keep up with the model technology, for all of us. I think it'll become increasingly more of a deeper infrastructure-level thing that software developers actually make sense of and figure out, for a given application, for a given use case, what's the right model technology to use. And hopefully for enterprises, for customers, the LLMs become a technology that they don't really think about. It's just something in the background.
Matt Turck [4:17] And you just mentioned open source, and I think you've said somewhere that 80% of the AI work may rely on open models. What's your sense of the reality in the enterprise today of the state of preparedness and adoption of open-source AI models today?
Arvind Jain [4:44] For most customer use cases, the companies are using the closed models at the moment, at least the ones that we are familiar with, as we go and work with large enterprises and we think about their core business processes that they want to automate with AI. Right now, the closed models are the ones that tend to dominate. And I think it's for the right reason. As you're in the early stages of the journey and you're trying to build an application where you're going to actually use AI as a key component, this is not the time for you to optimize and figure out what's the best model or what's the most cost-effective model because you're building an application and nobody uses it.
Arvind Jain [5:37] So it doesn't matter how much it costs. First, go and use the best, the state-of-the-art model, which today are still the closed models. And so, for that reason, you're seeing that largely people are starting with these larger models from OpenAI or Anthropic or Google. But as things reach scale, let's say you have an application that now has millions of users and hundreds of millions of daily interactions, now cost becomes a real factor. And that's when open-source models start to shine because you can actually take the LLM technology and make it work very specifically for your application, for your use case.
Arvind Jain [6:18] You can distill a model which is much smaller. It becomes not only a lot faster for you, it becomes hopefully a bit more accurate, but most importantly, it becomes much cheaper for you to run. And so that's sort of how we see the transition happen. Initially, you start your exercise, you work with the state of the art, the largest models, and then you start to fine-tune and distill to the right model as you start to get to scale.
Why enterprises choose open source AI (and when)
Matt Turck [6:52] And do you think that cost is the main driver for open-source adoption? As you just mentioned, is there a part of the reason why enterprises may want to use open source also something around the fact that they want to run it on their own infrastructure and possibly have a reluctance to send data out to commercial API companies? Is that something that you hear, or is that something that people make a big deal of, but in reality none of that is actually true?
Arvind Jain [7:15] Well, for certain businesses, there are actually real regulations and restrictions where you cannot send your data out. Having everything run in an air-gapped environment is actually a must. So for those, you have no choice but to work with open-source models. So that's one driving factor for it. Another one is cost. As I mentioned, when you have a large-scale application, you have to actually look at the needs that you have from LLMs and use very specific models for them, because they're going to be both better for user experience because they become faster for you.
Arvind Jain [8:07] They will hallucinate less for you because they are trained to do a very specific thing. So their range of errors actually also reduces. And then finally, they're more cost-effective. But there's another reason also, like the third reason that I would say why people actually want to work with open-source models is control. You feel, number one, you don't want to have dependency for your core business processes on other companies sometimes. And you want to actually, for that reason, run your own models.
Arvind Jain [8:25] You also feel like over time you'll be able to customize them better and better if you have the expertise to do that, that you won't be able to do if you were working with a closed model.
Matt Turck [8:54] All right, so more commercial models, more open-source models. And it's actually fascinating because there's even more companies being launched as foundation model companies, most notably Thinking Machines and SSI. I'm again curious, with your industry observer hat on, what you make of all of this. We were talking about those rumors, unconfirmed rumors, about a $7 billion valuation. And you guys are a fantastic business, reportedly well over now $100 million ARR. Meanwhile, you have companies that show up sort of out of nowhere and raise at staggering valuations: $10 billion valuation for one, $32 billion valuation for the other, before they even have a product publicly in market.
Matt Turck [9:11] How does somebody like you think about that state of things?
Arvind Jain [9:41] Well, the more LLM companies that are out there, the better it is for the industry. So we actually like seeing new names. We like seeing new models. And in fact, we benefit from them today in a big way. There's so many model choices that we have today compared to last year, models that are of comparable quality, models that are all getting better at certain things. For example, today, for a lot of coding use cases, people tend to use Claude a lot more.
Arvind Jain [10:11] For reasoning, they're using GPT. And so you're seeing these models get specialized to specific skills. And so, for application companies like ours, where we're trying to use this innovation and solve real business problems, the more innovation that we see, the better it is for us. And not just from a point of view that we have more technology to use, but also it creates that competition and it ensures that we can actually get the LLM advances in a more cost-effective manner that we can then actually bring our products in a more cost-efficient manner to our customers.
The agent era: what are AI agents and why now?
Matt Turck [10:59] Obviously, the big flavor of the day in the world of AI is agents. So we're going to unpack what that means in the specific context of Glean in a minute. But to start at a sort of high industry level, I'm curious about your view on the sort of state of agents in 2025. And I'd love to start with an actual definition of what an agent is, just to ground the conversation. And there's so many people talking about agents in so many different ways.
Matt Turck [11:16] So what is your definition of an agent? Well, I mean, I don't know if the world is looking for one more definition, but you had a great one on Twitter, which is what I'm riffing off of.
Arvind Jain [11:47] Well, I think for me, an agent is simply an application. It's a unit of work that used to be done by a human before. Now it's being done with the help of AI. So typically, any agent, like you think of, is going to work on some piece of information, some data, some knowledge, which can be brought from the public internet or the web, or it could actually be brought from within your company's systems. The agent works on this data. It has a task to perform.
Arvind Jain [12:19] It uses LLMs to actually reason and perform those tasks. There may be steps in this agent to actually self-reflect so that it can actually do the task better. And ultimately, it's going to produce some work, which is going to get saved in, again, maybe some of your internal enterprise systems. So an agent is nothing more than an application. We're all used to what software applications are. It's just that AI is being used in a fundamental way inside to do a lot of this work.
Automating business processes: real-world agent use cases
Arvind Jain [12:41] And so that's how we think of agents. They come in wide shapes or forms. Like today, everything from a simple workflow to a very deep, complex application, people are choosing to use the word agent to describe all of them.
Matt Turck [12:59] What's the why now of agents from your perspective? Why do they suddenly become such an important topic? And also, your sense of reality of what can be done today versus the hype of what people think they might be able to do in the near term.
Arvind Jain [13:33] So ultimately, an agent is automating a business process which was before managed by a human. And so that's sort of like, I'm talking in the context of agents in businesses. And of course, there are agents that you can be using in your personal lives as well. But for a moment, let's stick to business agents. So these are agents that are actually taking a specific business process and transforming them and automating them using AI. So, take some examples. You could actually, in your legal team, there's a process, like when you get third-party contracts, you have to redline them, make sure they conform to your company's guidelines.
Arvind Jain [14:14] So for that, today, you have somebody in the legal team, a lawyer who's actually reading these long documents, and they're going to redline it and pick all the places in the document where you need to actually change it to conform to your company's guidelines. And so this process, it requires a human. It requires them to actually go through a long document. They read it, they know what is the standard to which they're trying to actually align this document.
Arvind Jain [14:52] And you can now go and replace this whole process with an agent, an agent that actually automatically redlines your third-party contracts. And the way you build this agent is that you actually tell it exactly the process that human was following before, which is that, well, read the document, compare it with your version of the document, and then go and highlight or redline the problem areas in the document. And AI is pretty good at reading instructions, reading documents, and is able to actually go and now do this thing automatically for you.
Arvind Jain [15:33] And you can imagine that for many business processes today. Let's take a few more examples. Take customer service. One of the key activities there is customer service agents resolve tickets. And to resolve tickets, they're reading knowledge bases inside your company, looking for where the answers to these questions could be, and then use that to actually respond back to the customer. Again, AI can actually come and do that work. It can search over all of that information, understand the question, and answer that question automatically for the end user.
Arvind Jain [16:11] So, you can start to imagine that all these business processes that you have, you can automate them with AI. And this has been a desire of the industry for a long time. Before the AI wave, you had robotic process automation. It was a big thing. Now you can think of AI agents as just a much more advanced version of RPA. And there's a lot of business value from it. You can actually double the efficiency of every person in your company with it.
Are we there yet? The reality of AI agents in 2025
Arvind Jain [16:46] And that's why there's so much excitement. And it's actually like, the thing that's happening is that when you think about AI in this specific way—that you're going to build agents and you're going to automate your business processes—now it's super appealing to you because you can see direct ROI impact from building these agents, which is what is causing all the excitement in the industry today.
Matt Turck [17:04] And do you think that we are there or very close from a technology perspective, or do you think something is missing for the full realization of the vision, whether that's open industry protocols, MCP-style kind of stuff? What do you think?
Arvind Jain [17:15] We are in very, very early stages of this journey. If you think about agents, theoretically, an AI agent is promising you to do any piece of work. In fact, that's how, when we talk about even our product, Glean, the way we describe Glean to our customers is that you can come to Glean and you can ask any question to it, or you can actually give it any task, and Glean will use all of the world's knowledge as well as all of your internal company data and knowledge to answer those questions or complete those tasks for you.
Arvind Jain [18:11] Now, when I say that, I'm actually making this assertion that I can do everything in the world. And that's not the reality. Today, agents are still quite basic, I would say. They all still need a significant amount of supervision. I would think of agents as more being, well, I think I have a piece of task, I still own it as a human, but an agent can potentially do it for me. But it'll do its work.
Arvind Jain [18:36] I'm going to still go and review it. We barely see any application where customers are running agents in a fully unattended, unsupervised setting. So it's early, but the impact is actually very clear. But remember this thing: even if human supervision is required, that doesn't make the agent worthless. That same example I was giving you before about a legal person reviewing a 100-page contract and redlining it, that was a two-week activity that if you actually get an AI agent to redline it for you, it's going to do it in one minute.
Arvind Jain [19:09] And it won't be perfect, but now I can go in and look at the work that AI did and potentially spend like two hours instead of a week to actually go and redline that. And that's more than 90% savings for me. So I think that's sort of where things are today, that there is real value being delivered to customers, but agents are still better run in a supervised manner where a human is in charge and looking at the work of the agent.
Glean’s origin story: reinventing enterprise search
Matt Turck [19:58] All right, I'd love now to go into Glean itself very specifically and perhaps rewind back to the beginning, which I believe was in 2019. You started as an enterprise search company for the first few years of the company, which I find to be a fascinating idea because there was this whole generation of companies in the 2000s and perhaps all the way to the early 2010s. And I'm thinking of Fast Search & Transfer that was out of Norway and acquired by Microsoft for like a billion-something.
Matt Turck [20:45] There was Endeca that was bought by Oracle for a billion. Famously, Autonomy, which was acquired by HP for $10 billion, and then HP wrote off $8 billion a few months later. And Verity, that whole generation. So it sort of feels like it was an industry that came and then went away in the early 2010s. So what was the vision and the insight then that the industry was ready for reinvention when you started?
Arvind Jain [21:14] I would say it was not so much a vision or a realization that this was the right time to solve a problem. It was just personal frustration that I had, which forced me to start this company. Finding information at work has become increasingly difficult over the years because, well, number one, there's so much knowledge that we have in our companies these days. We live in a data-driven world. Number two, that information has increasingly gotten fragmented as we went through this SaaS revolution.
Arvind Jain [21:46] Every business has ended up with hundreds of these cloud-based applications because buying applications is so simple now. So businesses will end up with hundreds or sometimes even 1,000-plus applications, and all of your company's data and knowledge is spread across all of those systems. And my company, before we started Glean, was one of those companies. We built in the modern SaaS world with tons and tons of information and data with 300 different applications, and nobody in our company could ever find anything.
Matt Turck [22:03] And that company famously was Rubrik, now a wonderful public company.
Arvind Jain [22:25] That's right. So it was a personal problem, pain that I had. And I did feel like this problem had to be solved because it was not unique to us. Every person who I talked to, they would all say the same thing: that it's hard to find information. It's hard to get the answer that you need to do your work. And even before Rubrik, I was used to at Google, and ironically, we were making it easy for everybody in the world to get answers to their questions, but not helping ourselves internally.
Arvind Jain [23:06] Inside Google also, it was really, really hard to find information. But in 2019, there were a few key technical trends which actually made this problem more tractable, or actually more acute. The same thing, like I talked about how SaaS fragmented your information so much, making it hard to find. But SaaS also allowed you to actually tap into enterprise data and knowledge more easily because we have open APIs. And so, if you're building a search engine, it's much easier to build it now because you can actually tap into these standard APIs to these products and actually get that enterprise content all in one place and then make it searchable.
Arvind Jain [23:54] But the second big trend that we thought was super powerful was transformers. So it's 2019, and nobody as such is really talking about transformers and language models. This was still largely within the domain of search. Like, in Google, people were using these models to try to actually advance Google Search. And our team, me included, most of us actually came from Google, and we were seeing this big impact that the transformer technology was having on search. Suddenly, you could move from that keyword-based word-matching technology to find the right information to conceptually understanding questions and knowledge and doing the matching at that semantic level.
Arvind Jain [24:18] So we saw a big opportunity and we decided to actually use that to build a really good enterprise search. So that's the origin of how Glean got started.
Matt Turck [24:31] So that was phase one, and then there was a phase two, I believe, around the ChatGPT moment when you went from AI-powered search to something that's more like RAG. Is that fair?
Arvind Jain [25:03] That's right. So these models kept getting better over the years. I think for us, using transformers in 2019 made us one of the first companies to bring LLMs to the enterprise, but we didn't necessarily have the foresight. I should correct myself. We didn't have the foresight, and we didn't know how fast these models would get better. In 2019, they were good at understanding content, but they're not good at generating. They're not good at reasoning.
Arvind Jain [25:36] And these capabilities started to come over the years. By 2023, we were seeing this amazing ability for these models to actually generate their own answers to questions that people have. And we felt that actually allowed us to really advance our product in a significant way. So we started out with being the Google for your workplace, and now we could actually become the ChatGPT for your work environment. When people come and ask questions in Glean, instead of just surfacing the right information back to them, which we think is most relevant, we could actually make AI read that information and generate its own precise answers so that you could save our users even more time.
Arvind Jain [26:12] So we saw that opportunity naturally as the models got better, and we transitioned and launched our Glean Assistant, the product that you can think of as a more powerful version of ChatGPT inside your company.
Matt Turck [26:36] And the timing was perfect, right? Because you had done all the piping before of integrating the data sources, building the connectors to the search engine. So by the time ChatGPT came out, then you could just leverage that whole infrastructure to be able to do RAG on AI models internally.
Arvind Jain [26:37] Exactly.
Glean agents: from apps to agentic platforms
Matt Turck [27:05] Fantastic. All right, so that was the second phase of the company. And then it feels like the third phase is just starting. So we just talked about agents. So on February 12th, just a few weeks ago, you launched Glean Agents that you called a step into the agentic era, which feels like step three of the master plan of Glean. I'd ask, in a context where, again, agents are sort of the flavor of the day and everybody and their brother is launching an agentic platform, what's your sense of Glean's right to win, quote-unquote, the often-used expression?
Arvind Jain [27:53] So our journey with agents started early last year when we launched what we used to call Glean Apps and Glean Advanced Prompts. The idea behind agents for us was that we had already built these deep integrations into all of our enterprise applications. We could read the knowledge inside those systems, but we could also take actions within those applications. And this was a core part of our assistant product. So when you come into Glean Assistant, you could actually say, hey, do I have the following benefit?
Arvind Jain [28:32] Can I actually avail of this benefit or that? But you could also actually do things. You could say, hey, can you file a PTO for me? And so we are seeing a lot of demand from our customers wanting to take actions and do some work while they're within that Glean conversational interface. And so we started to build support for actions into these enterprise applications. And so this is where we were. We still are only a one-product company. It's this conversational chatbot like ChatGPT, but that knows everything about your company.
Arvind Jain [29:12] But this platform that we built was actually quite powerful, and we started to get a lot of demand from our customers. They wanted to do some very specific things with it. They didn't want this general-purpose assistant that is connected to all of the world's knowledge and the company's knowledge. They wanted to say, look, I want to actually build an HR first responder. This HR first responder application actually does only a few specific things. It answers people's questions.
Arvind Jain [29:38] If they are on human resources-related topics, and it can actually resolve some of their tickets using automation. And so our customers were saying, look, allow me to build this one specific application, which does this very specific thing, and which can actually take some additional instruction from me on how to do this piece of work. And so we launched this concept of applications last year that allowed you to do exactly that: work on a specific set of knowledge, solve specific tasks using AI.
Arvind Jain [30:18] And then over the year, we got a lot of customers using our app platform. And now we come into 2025, and everybody's talking about agents. So the first thing that we did was, in February, we renamed our applications product. What we used to call AI applications are now called AI agents. So we don't confuse the market, we have to sort of adopt the terminology that is becoming more standard in the market.
Matt Turck [30:18] Yeah.
Arvind Jain [30:52] But agents are now getting a lot more powerful. They're shifting from basic two-step RAG kind of application flow, where you take a task, you find some information, and then you make AI work on it to generate the right artifact. You're moving from that to actually running very, very complex business processes. There could be a hundred-step business process that you want to automate with AI now. And so the new version of agents that people expect now in the industry are significantly more complicated.
Arvind Jain [31:11] They have things like evaluations, insights, self-reflection, and things like those built natively into the core agent functionality.
Horizontal vs. vertical: Glean’s strategic platform choice
Matt Turck [31:22] Great. And I read somewhere that you had automated 50 million-plus agent actions. That was through the apps. Is that what you're referring to? That's right.
Arvind Jain [31:22] Okay.
Matt Turck [31:40] And then, in terms of the design of the agentic platform, you went very horizontal, basically enabling anyone to build their own agents with Glean. Walk us through the thinking. You could have gone very vertical per function. How did you think about this?
Arvind Jain [32:11] Since we started as a search product, where our objective was to help people find any piece of information that they're looking for inside the company, we built this horizontal platform. We built integrations to hundreds of enterprise applications, and we look at data and information inside each one of those systems so that when somebody's looking for something, we are able to immediately point them to the right pieces of information. So we built the platform before knowing that it would be used to build AI agents.
Arvind Jain [32:53] But now that we have this platform, which is connected to your Salesforce and your SharePoint and Workday and Google Drive and all these different systems—if you think about agents fundamentally, agents are working on some of your enterprise information and applying AI to perform some business logic, and then maybe saving the resulting work that the agent did back into your enterprise applications. So we felt that agents in general all tend to have these common elements. And the Glean platform actually solves for all of those because now, whether you're building an agent for an HR team or finance or legal, all of those systems that those teams use are already connected to Glean.
Arvind Jain [33:25] So that allows us to actually help our customers build agents across these functional use cases. So it's sort of like the reason why we started with the horizontal approach was just timing, because the first product that we built was search. But it has proven to be a very popular thing with our customers because today, one of the challenges that enterprises feel with AI is that there's so many vendors, there's so many products out there, and how many new products is an enterprise going to buy?
Arvind Jain [33:59] And I think the fact that Glean is this one horizontal platform that can actually do a lot for you is appealing to them, that they're able to try out AI across their different functional teams with just one purchase.
How Glean’s enterprise search works
Matt Turck [34:22] Plus, as you were saying, they're already there, right? That's the beauty of your strategic positioning, is that you're already the place where people have a dialogue with an AI every day to find information. So why not just ask the AI to do more for them? Wonderful. All right, so that was the evolution of Glean from enterprise search to RAG to agents. For this next part of the conversation, I'd love to get a little more technical and sort of talk about Glean under the hood, if you will.
Matt Turck [35:03] So starting with enterprise search, you came originally from Google, which was obviously the world of PageRank. Enterprise search is a whole different beast. So how does your enterprise search work under the hood, between the AI models and other kinds of heuristics that you've built to get great results? Because I should preview this: by all accounts, or anybody that's used Glean Search, the product is incredible in terms of quality of results. So how does that work?
Arvind Jain [35:27] Maybe first, let's understand the overall architecture of the product. So the first thing that you have to do to build a search product is you have to get hold of the information that you want to actually search over, which means in an enterprise, you have to build integrations with all the systems that a business uses. So that's the first part of our Cortex stack. We build these connectors and deep integrations into products like Confluence, Jira, Workday, Google Drive, and Slack, and so on and so forth.
Arvind Jain [36:03] Now, once you have the data, you have to actually solve a whole bunch of problems. Number one, enterprise search is fundamentally different from web search in the sense that the information inside your business is protected. Not everybody can see every document inside the company, unlike on the web, where we can all see all the web pages that are out there. So when you build a search product, you have to understand an individual. You have to understand permissioning of content inside your company for any document that is out there or for any message that's been exchanged in Slack.
Arvind Jain [36:35] You need to know who are the people who have the rights to see that information. And that has to become part of your core search stack. You have to build a secure search experience where, knowing who the user is, you only surface information to them that they have permissions for. So that's a big change. Nobody had done that before in the enterprise. Search systems traditionally have been built in a model where you dump all the content to the search system and it makes it all searchable for everyone.
Arvind Jain [37:15] So we do actually make it more personalized. The third thing is that enterprise knowledge is actually quite complex. Think about a company that's been there for the last 50 years. They've got tons and tons of documents that have been written over the years, conversations that have happened. And most of them are actually obsolete now. They're no longer—in fact, they have wrong information in them for the most part because things have changed.
Matt Turck [37:15] Mm-hmm.
Arvind Jain [37:41] And not only old content versus new content, but even if you think about a lifecycle of any project inside a company, you'll write a document and then you'll write another one, then another one. And there'll be 20 revisions of the document before you get to the final blessed version of a PRD, as an example. You have people making copies of those documents for their personal purposes. So ultimately, you end up with this environment where, on any given topic, there's tons and tons of information and you don't know which one is the right one.
Arvind Jain [38:02] And that's where a good search product comes into play. And you have to figure out what content is high quality, what content is written by the right subject matter experts.
Matt Turck [38:08] Right. You have a concept of expert as well, right? So authority is partly derived by the source.
Arvind Jain [38:34] Exactly. Yeah. You have to look into those factors. You have to look into engagement. Typically, if there's a document that is viewed a lot by people, it's probably for a good reason. It's probably a high-quality document. It's probably still up to date. So you have to observe. You have to actually look into the enterprise. You have to look at how people are interacting with knowledge inside the company and use that information. Because ultimately, a good search product learns from humans.
Arvind Jain [38:51] Humans and human behavior. And so that's what we did. These are all the different pieces of challenges that you have to solve. And then language models play a big role in terms of doing that matching. Language models are good to, on any given question, especially if you customize and fine-tune embeddings for your own enterprise, it's a pretty effective way of taking any question, bringing the right information back, but then using all those traditional search techniques to figure out which information is the most recent one, which one is the highest-quality one, things like that.
Arvind Jain [39:18] So that's sort of roughly how the search stack works.
Staying LLM-agnostic: integrating new AI models
Matt Turck [39:42] Great. Yeah. So you wrote somewhere that the success of AI was a question not just of models, but systems. So that's exactly what you described, right? You built a retrieval system around AI models. You're also LLM-agnostic with different models. I think you talked recently about how you just integrated Gemini.
Arvind Jain [39:44] Yeah.
Matt Turck [40:08] I'm curious about how that works in practice, whether you would route certain queries to certain models. And as an addition to the question, a lot of people talk about being LLM-agnostic. I just wonder what that means in reality, because as you see those new models come out, they behave in a very different way than the prior ones. What does that mean for a company like Glean? Does that mean you need to just stop what you're doing, look at the new model, evaluate it, and figure out how to integrate it so that you don't end up delivering an unsettling user experience with very different behaviors?
Arvind Jain [40:44] Well, that's the reality of an AI company today, is you have to fundamentally learn how to work in an unstable environment. This technology is moving so fast. And so, yeah, you have to do that. As new models come, you don't have the luxury to not look at them. You have to look at them. You have to see the new capabilities. You have to actually build the right evaluation frameworks to quickly see, as a new model comes in, we need to have a way within 30 minutes to know how well it's going to do on our product.
Arvind Jain [41:25] I mean, and you can do those things. You can actually build the right evaluation frameworks, things like that. Now, the way we ship these models to our customers is, one, we just give them the choice. So as they build agents in Glean, or when they ask questions in Glean, the user can manually choose what LLM they actually want to use to solve that particular question or that task. As you're building these complex multi-step agents, for each step, you can actually choose what's the best model to use.
Matt Turck [41:35] Mm-hmm.
Arvind Jain [42:04] So part of it is just working with advanced users who can make that selection for themselves, but that's the minority of the people that are going to be using your products. So when it comes to individuals, we have to actually make those smart choices ourselves. Given the nature of the question, can we use a fast and small model? Because we think the complexity of the question is not as high and a small model will produce good enough. So we have to make those kinds of judgments to actually automatically do the routing or choosing of the right model.
The architecture of Glean agents: tool use and beyond
Matt Turck [42:36] And still on the architecture front, now talking about agents specifically, you could have chosen different kinds of agent archetypes. It could have been computer use, it could have been tool use, it could have been specialized agents, and you chose tool use. Can you walk us through the reasoning here?
Arvind Jain [42:55] Yeah, so I think those are all actually, by the way, techniques that you have to use, all of them, in reality. I think we don't have computer use support in our agent platform. That doesn't mean that we've not chosen it. That probably means that we haven't gotten to it. And we prioritize based on customer demand. So right now, when you look at the kind of agents that our enterprise customers want to build, most of them tend to have this pattern, which I mentioned before, that they want that agent to work on a certain amount of business information.
Arvind Jain [43:38] They want AI to sort of do some work on it, and then they ultimately want to go and save the work of that agent somewhere within your enterprise applications. So in this model, to actually fetch the right data, you use the tool-based architecture. You build individual tools to fetch information live from different systems. You also build tools to actually save information back into those systems. And so that's the natural architecture that, in fact, everybody uses.
Data flywheels and personalization in Glean
Arvind Jain [43:50] Any agent framework that you will see out there will have the support to actually use external tools.
Matt Turck [44:17] I'd love to talk about data flywheels, data network effects, and personalization, which I think was implied in certain things you said. So the fundamental benefit of AI, one of the key benefits of AI, is that systems are supposed to get more intelligent with usage. What does that mean at Glean in practice?
Arvind Jain [44:44] So, two things. One, we actually learn a lot from that connectivity that we have into all of these different systems. I'll give you a simple example. Let's say that somebody asks a question in Slack on some technical topic, and some other individual answers that question and attaches a link to a document that contains the answer. So you can actually learn from this interaction. There are a lot of things to learn from this interaction. Number one, you've learned that that particular document that was linked in the answer contains the answer to that particular question that the user asked.
Arvind Jain [45:15] And that association is going to be super helpful to you in the future when somebody asks a similar question. You're going to use this document with more confidence than you could. Second thing is that you also learn that the individual who answered that question is potentially an expert on this topic. And so you can now, when there are more questions on this topic, actually, if AI cannot answer the question, point to this person as the person to go to, or you could actually take content written by this individual and weigh that more heavily on that particular topic in your retrieval systems.
Arvind Jain [45:59] So you're constantly learning. You're constantly learning from human activity. In fact, that's the source of all intelligence within the Glean platform, is that human behavior. And as you see more and more of it over time, it just gets better. And us being horizontal allows us to learn at the maximum pace because we are looking at what people are doing across all of these different systems.
Matt Turck [46:11] And I saw that you also have a concept of AI Prompt Studio, where you collect the collective knowledge. That's personalization as well, presumably.
Arvind Jain [46:42] So in fact, the prompts—it's like we're combining prompts and apps into this new concept called agents. Yeah. But the idea is that, yes. And look, the interesting thing with AI is that nobody knows what these models can actually do for you. There's a lot of uncharted, sort of unexplored frontier when it comes to these models. But you as an individual, sometimes you make a cool discovery, and that discovery ends up being something really cool for you. If you had a really easy way to share that with the rest of your team, that's how you sort of get those learnings going.
Arvind Jain [47:00] And so that's what our prompt library basically allows you to do, is when somebody has a great result with AI, we allow them to share it with their teammates.
Moats, competition, and the future of work with AI agents
Matt Turck [47:21] So as we get close to the end of the time we have, you mentioned that there's a lot of agents, a lot of companies in the space. How do you think about moat at Glean and how you build an increasingly defensible business over time?
Arvind Jain [47:48] This is the question that our employees have been asking us quite a bit as the market heats up, as more competition comes our way. My belief with any startup is you don't think about what the moat is. You just have to work hard on solving user problems. So we work with our customers. They have so many demands for us, so many things that they want. And I think just working hard and solving those problems one at a time ultimately allows you to build a very robust tech stack.
Arvind Jain [48:32] And the faster you build, the bigger that technology stack becomes. And I think that's sort of what ultimately the moat is: the hard work over the years. In enterprises, it's sort of hard to explore. It's hard to actually have network effects like the kind that you have with eBay. Enterprise software doesn't work that way. So we fundamentally just believe in that: build a good product, move fast, and that's how you stay ahead of competition.
Matt Turck [48:55] So fast forward a few years. If Glean fulfills its mission, what does that look like? Do you end up with a central AI in a company which knows everything about everyone across departments? How do you think about how Glean impacts the future of work?
Arvind Jain [49:27] So first, I think we believe that in the future, AI is going to be delivered in a way where there's going to be a set of horizontal AI technologies that are sort of organization-wide inside a company, LLMs being one of them. All the LLM technology is being delivered in a horizontal way to your enterprise. Similarly, the search and the retrieval and overall AI data platform is going to be delivered horizontally. And that's a platform like Glean, which is connecting all of your enterprise knowledge in one place and making sense of it and making all of that knowledge available to AI in a safe and secure way with the right policy, governance, and security.
Arvind Jain [50:11] Now, on top of that, you'll have agents which will proliferate, and they're going to be like thousands of these agents inside a large enterprise, and they're all functional and they are vertical in nature. Some of those agents will be built by you as a customer. Some of them may be specific vertical agents that you may buy from third-party specialists. And that's the architecture. Now, our vision, the way we want to be relevant in this world, is number one, we want to be the world's best personal assistant.
Arvind Jain [50:49] For employees at work. Glean Assistant, we believe, already is that today. You can think of it as a much more relevant, much more powerful version of ChatGPT inside your work. And the future that I see with AI is that every person who works is gonna have this amazing team of assistants, coworkers, and coaches around them that is going to actually make them a lot more effective. And this team is going to be proactive and help them whenever the individual needs help.
Arvind Jain [51:19] And today you have these kinds of teams. The CEO of a company has the luxury to have that team, but in the future, everybody, regardless of their role in the company, will have such a great team around them. And that's Glean's vision, that we want to be that team of AI agents around every individual that helps you do your great work. And the fact that we are connected to all of your enterprise systems will allow us to actually continue to be that, the best personal agent.
Arvind Jain [51:38] And then on the other side, we actually serve as a horizontal layer to power functional AI applications inside your enterprise.
Matt Turck [51:45] Fantastic. Very inspiring. Well, that feels like a wonderful place to leave it. Arvind, thank you so much for spending time with us today.
Arvind Jain [51:48] It was a pleasure being here. Thank you so much, Matt.
Matt Turck [52:09] This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.