AI That Ends Busy Work — Hebbia CEO on “Agent Employees”

The MAD Podcast with Matt Turck · with George Sivulka, Founder and CEO, Hebbia

George Sivulka is the Founder and CEO at Hebbia. We cover why AI agents become nodes on the org chart, why general-purpose AI outperforms narrowly vertical tools, and how Hebbia’s Matrix platform combines preprocessed documents with recursive decomposition instead of relying on RAG alone.

Watch on YouTube

Chapters

  1. 1:46 — What is Hebbia
  2. 2:49 — Evolving Hebbia’s mission
  3. 4:45 — The founding story and Stanford's inspiration
  4. 9:45 — The rise of agent employees and AI in organizations
  5. 12:36 — The future of AI-powered work
  6. 15:17 — AI research trends
  7. 19:49 — Inside Matrix: Hebbia’s flagship AI platform
  8. 24:02 — Why Hebbia isn’t just another chatbot
  9. 28:27 — Moving beyond RAG: Hebbia’s unique architecture
  10. 34:10 — Tackling hallucinations in high-stakes AI
  11. 35:59 — Research culture and avoiding industry groupthink
  12. 39:40 — Innovating go-to-market and enterprise sales
  13. 41:57 — Real-world value: Cost savings and new revenue
  14. 43:49 — How AI is changing junior roles
  15. 45:55 — Leadership and perspective as a young founder
  16. 47:16 — Hebbia’s roadmap: Success in the next 3 years

Transcript

What is Hebbia

Matt Turck [1:48] Good to see you.

George Sivulka [1:55] It was, I think, almost two years ago to this day that we started talking about Matrix in a very similar setting.

Matt Turck [2:00] Actually, Matrix was, if I remember correctly, not even a name then. It was stealth.

George Sivulka [2:10] It was pre-Matrix lingo, but it was still in stealth. A lot of the early ideas—and you were my first podcast ever—so I'm very excited to be back.

Matt Turck [2:22] Wonderful. Well, welcome back. So maybe for anybody that did not attend that event or does not yet know about Hebbia, what is the 45-second elevator pitch for what you do and why it matters?

George Sivulka [2:44] If we're in an elevator, Hebbia is—we actually build AI agents for finance, for investors, bankers, and lawyers. So anyone that does white-collar work, whether you're doing due diligence or discovery as part of what you work on, we build an AI agent platform so that you can leverage some of the latest and greatest.

Evolving Hebbia’s mission

Matt Turck [3:05] And rewinding back to that chat from a couple of years ago, I remember that you had a very powerful kind of mission statement, which was to keep smart people from doing stupid tasks. And if you watch that back and if you sort of fast-forward to today, are you still doing exactly that? In what way has the mission evolved?

George Sivulka [3:35] I think when we started Hebbia, one of the fundamental insights was that you had a lot of the smartest people in the world doing the stupidest tasks, and that it seemed like there was an arbitrage opportunity there to save them some time or some effort or some sweat. The goal of Hebbia was never to just stop at really highly paid knowledge professionals, the investor or the finance guy or the lawyer. It was actually always to go much broader. And our vision and mission have solidified really over the last couple of years into building a capable AI platform for a billion people.

George Sivulka [4:10] And the idea is that there's lots of consumer AI products where you can go and talk about your sushi restaurant and wine in San Francisco, but there's actually not very horizontal, general-purpose workplace AI products. And what we mean when we talk about capable AI and AI that can do things is actually something much more horizontal. It's something that's much richer in its interaction. It's actually not as verticalized as you might think, but it still has the depth of what an expert platform would be.

George Sivulka [4:43] So if you think about how that vision and mission have evolved, we started on Wall Street, we started with lawyers, we started with investors and bankers, but now we've built an AI platform for experts that lets anyone, whether you can code or you don't know how to code, or you can only use Cursor to code, build whatever kind of knowledge work application or agent you'd like.

The founding story and Stanford's inspiration

Matt Turck [4:53] Could you rewind back to the initial sort of light bulb moment when, I guess, you were a student at Stanford, if I remember correctly? How did that all come about?

George Sivulka [5:17] It was 2020, and at the time, around June or May of 2020, no one was talking about AI. No one was talking about large language models. And I remember, I think the topic of what I was studying at the time was meta-learning, or the idea that you could figure out how to build a machine that could learn how to learn or generalize to any task. And I thought that meta-learning would become the most important software of all time. And at the same time, I was working on it in my research.

George Sivulka [5:38] And then all of a sudden, June of 2020 comes around, and OpenAI actually released GPT-3. So before ChatGPT, I remember using it, and I think after using it a few times, I sat back in my chair. I was like, this thing is a meta-learner. It has beaten me to the research punch that I was working on. And if I couldn't actually play a part in creating that fundamental technology or that technological revolution, I knew I wanted to be a part of applying that in a really interesting and meaningful way.

George Sivulka [5:58] And I made the bet that, hey, this would be worth basically throwing my whole life, my whole research career away on and starting a company from scratch.

Matt Turck [6:07] And not to make you blush, but so you were doing, like, a PhD. You were like 21 or 22 when you were starting your PhD.

George Sivulka [6:11] I'm blushing. Stop. Okay, not really.

Matt Turck [6:22] Stop telling people. Stop forcing me to tell people how smart I am. No, but there was something like that, right? Like, you graduated super early from Stanford. When did you get your PhD? Just tell that story in two sentences.

George Sivulka [6:40] I was a young hustler, is the reality of the story. But I was one of the youngest people to graduate Stanford and was the youngest person in my PhD program that was fully funded. So I was throwing away, at the time, millions of dollars of research grants and not having to be a TA for the extent of my graduate school, which was actually quite the luxury. Didn't know I could raise, didn't know we'd be able to build the company that we built and the team that we built.

Matt Turck [7:04] Great. So speaking of that, what's happened over the last couple of years? Any metric, including vanity metrics, that you can share? Fundraising history, number of customers, number of documents, whatever it is that you want to share to give people a sense for the reality of the company as of today?

George Sivulka [7:29] I can't forget that you're a VC, so I'm not allowed to share any metrics on here, but there's some public ones. And I think things that I'm really proud of, really lasting, are first and foremost our team. So we built a team of 100 amazing, incredibly smart folks, all five days a week, sometimes six, in office in New York City. And we are now just starting to become a multinational corporation. And so we're opening San Francisco, and we've already opened a London office, which is incredibly exciting, with goals to end the year at 300 to 400 employees.

George Sivulka [8:01] Lots of exciting growth, a lot of it here in Silicon Alley and, unfortunately, Silicon Valley. We have to move out there a little bit. But I think on AI metrics, or a little bit more about the product and how it's used, one of our favorite things to track is the amount of unstructured data or the amount of pages that are processed by the platform. And a really interesting thing is, hey, last year, Hebbia and probably all of the other major consumer model providers processed around 100 million pages, probably around, whatever, hundreds of years, maybe thousands of years of reading.

George Sivulka [8:45] This year, we're already on track to process around 4 to 5 billion pages. So somewhere around 50,000 years of reading for a human that's taking the right amount of breaks. And that exponential is just one of the most phenomenal curves that you've ever seen. It's like in a big board in our office. And I think the reason we're so proud of it is, whereas other AI platforms, you're having a bit of a transactional relationship with the AI. You ask a question, you get a response.

George Sivulka [9:14] With Hebbia, you can give it these complex tasks and it really churns through vast quantities of data and does work the way you work. It's much more of an agent than a chatbot. And so when you're actually looking at the work that it's doing, or that it would take in people-hours, a team of really highly paid professionals to do, it's actually having a massive impact at the organizations we've rolled out. Now we're deployed at, I think, between 40 and 50% of the world's largest asset managers, some of the tier-one investment banks in the world.

George Sivulka [9:37] We have an incredibly fast-growing legal segment. So I think even just the share of our revenue that's in law is increasing at an exponential clip. And so it's just an incredibly exciting time to be at Hebbia and in AI.

The rise of agent employees and AI in organizations

Matt Turck [9:51] So we're going to go into all things Hebbia, the product, the tech, and all the things. But before we do that, just maybe six or seven days ago, you had a really interesting blog post that you published where you talked about the concept of agent employees.

George Sivulka [9:53] Yes.

Matt Turck [9:58] So, going to that, what is an agent employee in your world?

George Sivulka [10:21] I think that the idea behind the blog post and talking about an agent employee actually stemmed from a way that I think organizational design is starting to change at our customers. And I actually likened it to an old organizational trichotomy of remote work, where the internet actually came out and took—it basically decoupled the output of labor from where you were. So you could be all in the same room in New York, or you could be across the world, and the internet basically democratized access to talent.

George Sivulka [10:56] With AI and these agent employees, you're actually also starting to see, hey, I'm going to have full-time things, whether they're employees or not employees or whatever you want to call them, where instead of actually having output or work or labor coupled with personhood, it will now be completely decoupled, not only from location, but from whether or not you can pay salaries or have humans doing the jobs. So the entire essence of the piece was, hey, just as we have remote employees, we have hybrid organizations, fully remote organizations, and in-person organizations, you're going to start to have fully human organizations, actually fully AI organizations, like the one-person billion-dollar startup or these things that are effectively just APIs.

George Sivulka [11:46] And you'll actually have—and this will be the most common thing—hybrid AI and human employees, agent and human employees working alongside each other. And when you think of that from an organizational design perspective, you can start to define the agent employee as another node in your org chart. And it becomes very clear that if that's a node in your org chart, it will probably have to have an email. It'll probably have to have a Slack. It'll probably be doing things the wrong way, and you'll have to manage it in the right direction.

George Sivulka [12:20] And the makeup, whether you're 50% agents or 10% agents or 90% agents or 0% agents, will actually define the output of the organization. Amazon is a perfect example of how this happened, how org design shaped what they built. If you look at AWS's offerings, every single offering in that big menu is a different startup altogether with its own GM. And that intentional organizational decision impacts their product. It impacts how it works. It impacts whether or not those products work together really elegantly.

The future of AI-powered work

George Sivulka [12:36] The same will be true as you start to think through these agent employees as nodes and how you organize them, and whether or not you're where you are on the 0 to 100% AI agent dialectic.

Matt Turck [12:54] So fast forward a few years. What does that all look like? Whether you call it 2030, let's say, not 20 years out, 2030. So we're all what, AI managers? We spend half our days interacting with agents. Do we still interact with humans? What does that look like?

George Sivulka [13:20] I think that right now, we're already AI managers. The only difference is the AI takes a single step. So if you are an AI-modern organization, you probably are using some sort of chat or RAG application, and how good you are at prompting that AI is how good you are at managing that AI. And instead of letting the AI go and prosecute a task over and over and over and self-correct, it can only take a single step. As agents are rolled out, you'll actually start to see people that are really good at prompting, really good at defining a process, be the best managers and actually be the best at extending whatever their agenda is in the organization or making their function the best.

George Sivulka [13:44] And so I think that everyone will be prompting, and prompting is managing, and it will all blur pretty soon.

Matt Turck [14:10] As a quick reality check, I think you mentioned somewhere that AI agents will contribute more to GDP than human workers within a decade. How backloaded is that? In other words, what's your sense, as somebody who's in the proverbial trenches every day, of the reality of agents in the enterprise in terms of what they can do and what's overhyped?

George Sivulka [14:37] We will still be deploying AI in its current form as a chatbot in 10 years. I think the way to think about that is there's actually still plenty of businesses in New York that don't take credit cards. They only take cash, and organizational change and technological change will take a lot of time and a lot of effort. At the same time, most people use credit cards. And when things work and they're not just experiments, they actually happen very fast. You're starting to see that with certain AI companies that are massively penetrating markets, that have gone from zero to some double-digit percentage of their SAM.

George Sivulka [15:02] And we're fortunate that Hebbia is one of them. But at the end of the day, the stragglers, the long tail of adopters, will still take time. And I think that, to answer your question on the nose, everything is going to be backloaded. I think it'll happen in the decade, like the change to credit cards happened over five to 10 years, and now this change to just point and click in credit cards happened even faster than that. But there will still be stragglers.

George Sivulka [15:17] It'll still take time.

Matt Turck [15:42] One last question then before we get into Hebbia in detail. What do you find particularly interesting in AI today? AI research, open-source reasoning models that you may or may not sort of import into Hebbia directly, but what do you find interesting in terms of what's happening in research? There seems to be a new model every other day, a new thing every other day. What catches your attention?

George Sivulka [16:06] It's a good question. There's the felt sense, I think, in communities like this one and then maybe the larger AI community, that the scaling laws for training have slowed down a little bit. And we really haven't had a massive paradigm shift that has been released recently. At the same time, there's a lot of really interesting research directions in scaling laws for scaling during inference. And what that means is, hey, maybe we have a fundamental unit of compute, and that is a single inference, a single forward pass on a large language model.

George Sivulka [16:42] How do we actually now use that, or run lots of inference like an agent, like OpenAI o3, o4, and all of these more agentic, more reasoning models, to actually get better at doing these difficult tasks? Actually, the whole scaling-at-inference paradigm was pioneered at Hebbia. Our early Matrix product two years ago, we were one of the first people to say, hey, you get way better accuracy from using more large language model calls at runtime. And so we built lots of infrastructure to scale that up.

George Sivulka [17:11] We actually built an agents team before it was even called agents to actually go out and run these larger jobs. And I think that's probably the most interesting research direction moving forward, is like, hey, let's go and say this current scaling law of training has slowed down. The scaling law of inference, or test-time compute, is very interesting. Let's double-click there. There's lots going on with multimodal research. Obviously, context windows there, how you apply them, how you tie them into the larger ecosystem of foundation models that are used for text applications, is very interesting.

George Sivulka [17:45] And I think the number one problem in all of AI is elongating the context window. And so we have a few different metrics. It's now a focus. You've got these needle-in-the-haystack tests that measure how good they are. One of the areas that Hebbia has actually been doing a lot of research in has been artificially extending the context window, or effectively extending it. And that's been a massive boon, leveraging inference-time compute for accuracy for research and diligence workflows.

Matt Turck [18:01] Actually, since we're on the topic, I cannot resist asking for a double-click on this. So what does that mean? How do you artificially extend a context window?

George Sivulka [18:25] Well, the idea of a context window is that you can think about all of the things in the context window. So as humans, we currently can remember whatever, the last 10 minutes of conversation, maybe not this entire conversation, but we also have bits and pieces over our life and our training and our early careers that we've collected that make us really good employees. AI, you've got a system prompt, and then whatever, up to a million tokens to jam as much context as you can in there.

George Sivulka [19:01] When you think about what you actually want AI to do, you want it to reason over all of your data. Applications like RAG can jam the context window with as many search results as you want, but it's not going to reason over that data. It'll just search for stuff that exists. By elongating the context window, ideally, you'd be able to connect the dots, find stuff that's there, but also stuff that's not there, stuff that's missing. And what that means actually, tangibly, from a product perspective, is how do we very elegantly use the current context window, scale that up, right?

George Sivulka [19:29] Or jam as much of a current maximum context window into a single question to get the right answer. And so it's actually more of an infrastructure problem. It's more of a patchwork or convolution problem of how you actually apply that with the current length of context windows. And then it's actually kind of an information theory thing where you have to think through how what I get out of each of these runs can be maximized every time, but also maximally compressive so that every single future iteration gets just the right information, nothing that's not relevant.

Inside Matrix: Hebbia’s flagship AI platform

Matt Turck [20:09] As you mentioned, a year or two ago, you released Matrix, which is the flagship product for Hebbia. So let's get into it. What does that actually do? I think it was billed as the interface to AGI. What's sort of my reality as a user of the product? What do I have in front of me? What does it do for me?

George Sivulka [20:37] I think the most important job in the future will be how you actually manage AI agents or how you prompt these things at scale. You can think of Matrix as actually running a bunch of sub-agents, or an interface like a Trello board where you can assign a lot of tasks and then a bunch of agents will do these things. And so it could be a network or a multi-agent platform where you can basically say, "Hey, I want all of these things done." They all are related to each other in some specific way.

George Sivulka [20:47] And every cell in what looks like a grid output ends up becoming a task that's completed by an agent.

Matt Turck [21:01] And you mentioned you started in finance and then you expanded to law and consulting. How do you customize something like this? And relatedly, how do you even decide which next vertical you need to go into?

George Sivulka [21:29] One of the things that is a massive fallacy in AI applications today is treating verticalization as paramount. All of the VCs, a lot of entrepreneurs as well, believe, because it's been true for the last 20 years, that building a very verticalized piece of software is the only answer. And so you've got plenty of startups that are like, "Okay, I'm only going to be AI for blank, AI for compliance, or AI for law, or AI for X, Y, or Z." When you have an AI capabilities approach to the problem, you actually start to realize that generalization will beat specialization every single time.

George Sivulka [22:09] And what I mean by that is, let's say you want an investing agent. Your investing agent will become a researching agent, and the researching agent should become a learning agent. And then you get to something closer to AGI. If you want a legal agent, your legal agent will become maybe like a diligence agent or something that's very hyper-logical. And if it's really logical, it'll be basically the smartest possible agent, and it'll end up at AGI. Or if you're a banker, your banking agent will become a marketing agent because banking is effectively dressing up a company to be sold.

George Sivulka [22:37] And then the marketing agent will become the persuasion agent, and the persuasion agent will become something closer to AGI. And what I'm trying to highlight here is that it's a way to hack, to prompt one or two things for a specific set of customers at this point in time today. But as AGI comes closer or is actually here, as these models improve, and as we as people get better at using them, we're going to get way more interested in actually extending the context window or pulling these expert knobs or taking something that's general and writing our own prompts and building our own agents.

George Sivulka [23:27] Because the reason why you as an investor beat the market isn't because you're all using the same ChatGPT wrapper prompt. And the reason why you as a lawyer are getting paid $2,000 an hour isn't because, "Oh, I have the same shortcuts from an AI-for-law company." It's because you are an expert and you can customize the greatest and the best tools to the way that your firm works, to the way that you actually can find an edge. And that's not specialization. The best investors aren't only looking at SEC filings.

George Sivulka [23:55] They're reading really crazy data sources that have nothing to do with finance. The best lawyers, whether they're persuasive or logical, they have backgrounds that are way outside of law. You don't want specialization; you want generalization. And a lot of our customers are realizing that. That's actually the superpower of Hebbia. And that's why so many people are buying us.

Matt Turck [23:59] You don't seem to be a huge fan of the chatbot as an interface.

George Sivulka [24:01] Nothing against chatbots.

Why Hebbia isn’t just another chatbot

Matt Turck [24:09] Talk about how different the Hebbia interface is and what you're trying to achieve there.

George Sivulka [24:33] I think chatbots are, in my eyes, like the TI-84 or like the HP-12C, like a calculator. It's a one-off. I'm going to go and put in my equation and then get a response. Nobody does their taxes in a calculator. Nobody computes a DCF or does serious knowledge work in a TI-84. You stop using them in high school. You might have a simulator on your laptop because it's novel. The way that we work is we have these really broad, flexible interfaces called spreadsheets.

George Sivulka [25:09] And the way that everyone interfaces with a computer, or the fundamental paradigm of computing, is through a spreadsheet. It's flexible. It's highly modular. You could call it an expert system. And that's actually the way that billions of people use capable computing today. The spreadsheet in Excel is the most important software ever invented. Hebbia takes an approach that's not dissimilar. Just like a single cell in Excel is a calculator, a single cell in Hebbia could be a chat or it could be a variety of other AI operations. But the Matrix platform, how people use it, is so much broader.

George Sivulka [25:27] It's so much more flexible. It's an expert platform, and we're building that expert platform for how humans will interface with AI agents.

Matt Turck [25:34] And so the way they interact with the product—and realizing it's difficult to talk about an interface verbally—but it's more of a—

George Sivulka [25:37] You could have invited me for a demo. I would have done a demo for you guys.

Matt Turck [25:43] Next time. But it's a spreadsheet, right? I mean, it's—

George Sivulka [25:58] One of the modes. So we think we're building effectively the new Microsoft Office suite. And just as Microsoft Office had Word and PowerPoint and Excel, we have Matrix and then actually a variety of other applications, some of which I can talk about, some of which I can't. And we have a chatbot. And we think that there's going to be an entirely new productivity suite for how humans, how experts want to interface with AI without having to get in the weeds and code.

George Sivulka [26:09] And it's one of those applications. Yeah.

Matt Turck [26:27] Great. All right. If that's okay with you, a little bit of a technical deep dive into how that all works behind the scenes. So presumably, to start with, you need to connect with a bunch of sources of enterprise data and process the data. So how does that stage of ingestion work?

George Sivulka [26:49] Ingestion and indexing is one of the fundamental pieces of having any knowledge work application. There's many different steps to it. So first is just collecting the data sources and hooking into as many providers as possible. And that's not fun or technical. The more interesting thing is actually, once you have the data, how you index it. We believe that you can do a lot of preprocessing, and I'm not talking about keyword search and building a BM25 index or having some sort of semantic search index with embeddings, even your super-long embeddings that are coming out.

George Sivulka [27:18] That's not actually that interesting to our users. Again, that's good at doing searches. It's good at finding things in the data. But you want to start to process information before the user even asks a question. You want these agents to be doing work ahead of time. What we do instead is we actually ingest documents and preprocess, prepopulate, depending on the doc type, depending on the context of the document, a really rich schema and understanding of each document that ideally would be, hey, we've pre-indexed and pre-done 90% of the work.

George Sivulka [27:51] At the same time, that's not enough for asking and running user questions on the fly. This war just broke out here, or this crisis is happening here, or this event that you couldn't have predicted is happening. We use the same indexing engine and the same infrastructure to run things on the fly. And so you can consider any query that goes through Hebbia or that goes through Matrix as a combination of a lot of pre-indexed work and then some stuff that's on the fly and custom.

George Sivulka [28:25] And the pipeline to do that, from parsing documents, sometimes using multimodal documents to figure out formatting and structure and images, actually how we then go in and run that index and what we're retrieving over, and how do we save the work of the pre-indexing agents—all of those are interesting search problems that I could talk about for a while. That's a bit of the thesis behind how we index.

Moving beyond RAG: Hebbia’s unique architecture

Matt Turck [28:53] And then perhaps that's part of what you just described, but I wrote down something about what you call the ISD architecture. So that's the RAG part after that. And you were saying somewhere that RAG sort of doesn't cut it for very hard problems. So what is your approach to RAG? And I assume a lot of people here know, but maybe use the opportunity to define what RAG is in the first place.

George Sivulka [29:18] Sure. RAG is an architecture for using AI. That's retrieval-augmented generation. It was first coined in a paper in March of 2020 by a bunch of Facebook researchers. Hebbia was actually the first to turn that into a product. So it's a very close thing to my heart. Back in 2020, we were the first people to actually productionize it, roll it out. And the idea is that you could hook a search engine up to an LLM. It's basically what Perplexity uses.

George Sivulka [29:31] It's like a search engine to LLM, except over the web, you always have an answer. Over offline and unstructured and private documents, you don't always have an answer. And what we say when we say RAG doesn't cut it, or that Hebbia, actually, after making that as our baby, we had to kill the baby and we turned away from RAG to what we call ISD, is we said, well, we need an architecture that's closer to extending the context window versus just searching or calling an external tool.

George Sivulka [30:09] And ISD leverages that index, so a lot of work that has been done ahead of time, but then it also leverages a way of recursively reading subdocuments. So still leveraging tools, sometimes leveraging that infinite effective context window to bubble up an answer. And it's really a decomposition agent at its core. So you ask a question, instead of it just pinging a search tool, it can ping a variety of different tools, does rich decomposition, shows its work in the Matrix, and then synthesizes a complete response towards the end.

Matt Turck [30:32] Do you have pieces still of, I guess, almost classical RAG architecture, like re-rankers and that type of stuff somewhere? Or completely—

George Sivulka [30:57] Very funny thing, and my first five employees will get a laugh out of this and no one else, but we have, to this day, the most accurate re-ranker that has ever been released, and we do not use it. We spent a lot of the first year, year and a half, just training embedding algorithms, training different versions of ColBERT with multi-embedding architectures for a single passage, and then training re-rankers. And we came up with a novel re-ranker architecture, which four years later, academia and industry have not beat, and we do not use it.

George Sivulka [31:23] And the reason we don't use it is because a lot of the time, search isn't what's important. If we were building a public web API LLM, re-rankers would be important. But we scrapped all of that for a really heavy infrastructure play that uses tons of large language model calls. It is really expensive from a latency and just latent dollar-cost perspective, but achieves really accurate answers for deep research, for deep diligence tasks that people can only do on the heavier platform.

Matt Turck [31:48] Great. Let's talk about the model part. So I think you guys are very close to OpenAI. So do you use multiple OpenAI models? Do you use all the stuff that you can talk about, open source, what have you?

George Sivulka [32:16] Yeah, we've got partnerships with OpenAI, but also Anthropic and also Amazon. So we're playing the field. Don't tell Sam Altman. But the idea behind Hebbia is we believe the model layer will become commoditized. It's no longer a hot take. So whatever models you want to use, and hopefully eventually computing on the edge will be the way that you run these things, just like you don't ping the cloud to run your Excel model unless you're in Google Sheets.

Matt Turck [32:29] And then you've built some very smart stuff, from what I can tell, around sort of scaling all of this. So you have this concept of Maximizer, which sounds like it's a router, but smarter. What is it?

George Sivulka [32:59] Yeah, I think that some of the most interesting stuff at Hebbia isn't actually only the agents research, but actually the pure systems problems where running—I think we run around 250 billion large language model calls a month. And our workflows, I think before we had Maximizer, with all the rate limits that we had, we could run a million tokens a minute, and now we can do 500 or 450 million tokens a minute. All of that actually wasn't an AI problem. It was a systems problem.

George Sivulka [33:31] Like you mentioned, it is a router. You could think of it as a two-sided marketplace, but one of our systems there is called Maximizer. We liken it, from an analogous perspective, to an air traffic controller where, just like an air traffic controller basically tells you who can fly, where, and when—and apparently you're not supposed to fly out of Newark right now. I hope it's fixed—Maximizer basically has a handshake between lease grants and lease requests. So if a user or a job has a really high or low priority and wants to go and use GPT-4o, it can make a lease request, and a lease grant will say, okay, you have to use Azure or you have to use X, Y, or Z models.

Tackling hallucinations in high-stakes AI

George Sivulka [34:10] And it actually ends up with the theoretical information-theoretic maximum utilization of any rate limits. So not only are we running the most amount of pages, the most amount of tokens through our platform, but we're doing that in the most efficient, perfect way. And that's all our infrastructure, which is cool. I could get into more, but yeah, I don't want to bore everyone.

Matt Turck [34:26] Inevitable question about hallucinations. High-stakes professional context where people are paid a lot of money to provide very reliable results to the customers. Is that something that in 2025 is as much of a problem as it was?

George Sivulka [34:51] I think it's old news. And I think the only reason we still talk about it, it's like everyone talks about hallucinations and no one knows if it's happening or what's going on. It's like fugazi, fugazi. It was a problem back when the models were stupid, but they're obviously not stupid now. And I'd even say that they're way better than any human. So when we start to think about hallucinations, it's like, well, where does that come up? Have we actually seen a lot of that happen?

George Sivulka [35:17] Probably not recently. But the place where it does come up is actually a limitation of RAG. It's when you've got the wrong documents, where you're searching for something and it finds something that's broken. And what you're realizing is that the models are actually pretty good at doing reasoning. They're really good meta-learners. The hallucination, whether it's a problem or not, no one really cares. What people care about is when these things fail because they don't have the right context or they're not reading the documents in the right way.

George Sivulka [35:47] And that's why we've done a lot of research on how to feed the right stuff at the right time. And if that's really, really expensive, we don't care, because models are going to become—intelligence will become too cheap to meter. That's what our industry is predicated upon. So yeah, we'll spend $10,000 in model costs per user a year. We don't care. We don't, but indicative.

Matt Turck [35:57] You just mentioned research. How does that work at Hebbia? Do you guys have a concept of an AI lab within the company, or is it distributed across the team? And how important is research, especially in a context where, again, there seems to be something new every other day? How do you think about doing your own research versus just, like, free-riding? Yeah, it feels like there's something new every day, but also if you've been tracking the field, it also doesn't.

Research culture and avoiding industry groupthink

George Sivulka [36:43] It kind of doesn't feel that different. I don't know if I'm allowed to say that, but it doesn't for me. There's new models, but they don't feel that different. Some of them are more sycophantic. That's the news. I think not a lot of people are saying it, but I think SF has a really big groupthink where it's like, okay, everyone's going to go out and work on the same things. And there's this sort of memetic attraction rivalry between OpenAI and Anthropic, and what are they going to do next?

George Sivulka [37:13] And like, oh, they have a new interface, or the chat bar got a little bit jumpier with the animation. And we kind of know, like, okay, they can make videos and it's too slow, and they can make images and it's better now. The issue is to come up with new interfaces and AI that can really do new things. I think you kind of have to avoid the groupthink. And so we currently don't have any AI researchers hired out in San Francisco.

George Sivulka [37:42] We don't even listen a lot to what people are requesting and really try to think about, okay, what is the tool that AGI would use? Like, if you actually had AI that could choose any other AI tool to use, what would that look like? And we have that guide what the interfaces look like. That's kind of how we came up with Matrix and a lot of the other interfaces that we're talking about. That's how we came up with ISD versus RAG and what the rest of the industry uses.

George Sivulka [38:00] It's kind of by avoiding what is going on on Twitter and just keeping our heads down. So we do have a research lab. It is not a standard research lab. It does spend a lot of time on information retrieval.

Matt Turck [38:20] So do you think that, to take the extreme of what you said, that research is actually slowing down? So there's test-time compute that we were talking about that sort of felt like the major innovation of the last six months. But after we're done with test-time compute, are we sort of running out of tricks?

George Sivulka [38:44] I think there will be more tricks. I do believe that AI will start to work on itself. I don't mean to be relatively pessimistic, but I wouldn't start a company in AI right now. I mean that from the perspective of the alpha is gone. When I was starting the company, everyone was working on crypto, and it was like the alpha was gone. I was like, okay, we're going to work on AI. And so I don't know what the next thing is, but I do think that the models will continue to get better.

George Sivulka [39:19] They'll get way cheaper, they'll get way faster, and the user experience will improve. I think longer term, just as I'm talking about how generalization beats specialization, you saw over the last 20 years of enterprise SaaS Excel get unbundled into a million specialized things. It only made Excel more powerful, but there will still be specialization as a second wave to create a lot of value. But I don't see lots of very large changes in the near future. But maybe I've also gotten spoiled from how much stuff has changed.

George Sivulka [39:25] So I'll have to think about that one.

Matt Turck [39:33] Great. So everyone pivot to crypto right now. Announcing the rebranding of the event: Crypto Driven.

George Sivulka [39:36] Crypto, Crypto New York City.

Matt Turck [39:36] Yeah, exactly.

George Sivulka [39:38] There we go. I wouldn't come.

Innovating go-to-market and enterprise sales

Matt Turck [39:57] Switching tacks a little bit and talking about go-to-market. You've innovated there as well. I think, among other things, you have a very different take on who salespeople should be for a company like Hebbia in terms of backgrounds. Do you want to talk to that?

George Sivulka [40:05] Yeah, I'm happy to. I don't know if this audience is interested in sales playbooks and lineages and all that stuff, but—

Matt Turck [40:08] Yeah, I think there's a fair amount of people building companies.

George Sivulka [40:36] So, one of the things that I commonly think about is the best enterprise SaaS selling organizations over the last, whatever, two, three decades all come from the same lineage or the same DNA. So they all come from BMC and AppDynamics, and then they went to MongoDB, and they're all kind of like the same thing. They have the same playbook. They have a framework called MEDDIC or MEDDPICC, which is all around how you sell value. AI is very different from those organizations' products.

George Sivulka [41:08] If you actually try to apply this standard, okay, we're going to do this type of discovery and we're going to do this exact sort of MEDDIC sale, there's some things that will take and that you can extend to AI, but right now it's less around pain. It's less around standard enterprise SaaS cycles and probably more around FOMO, around missed upside, around value cases that are really hard to define. And so the profile of a traditional seller that would excel at these best-in-class playbook organizations—we'll see if that takes for Hebbia.

George Sivulka [41:42] If that takes for our peers in B2B AI application spaces, the jury's still out. At the same time, one of the things that we've seen really work are people with domain expertise, people that can speak to the customer's lingo, vernacular, processes, and work. And honestly, people that are very similar to consultants and kind of having that playbook consulting muscle together is, I believe, what will work when there's a true paradigm shift over the last 20 years of enterprise SaaS. It's really just unbundling of things that people already know exist.

Real-world value: Cost savings and new revenue

Matt Turck [42:04] So interesting, right? To that concept of FOMO buying. Do people still ask you about ROI, how to justify your price? And if so, how do you do it?

George Sivulka [42:26] Part of the beauty of Hebbia is that I think we're one of the only AI companies that builds value cases. And that is really significantly upselling our customers. We treat it like a traditional playbook company. We actually build a really strong understanding of if you're spending this much, this is what you're getting in return, or these are the expenses that you can actually remove. And we commonly partner with the world's best financial firms, not to just give them an AI tool like our competitors, but to actually go out and understand and say, okay, it's board-level. You're trying to impact your P&L to the tune of $100 million.

George Sivulka [42:47] These are the ways that we've seen it done. And it's like a bit of technology and a bit of consulting. Yeah.

Matt Turck [42:56] And is your ROI based on cost saving, meaning hours not spent doing a task, or is it based on increased revenue?

George Sivulka [43:16] There's cost saving. It's not only hours spent. Sometimes it's removing third-party legal expenses or third-party consulting expenses or third-party expert network spend. But a lot of what really is driving value and capturing the imagination of customers is the idea of, hey, you can actually make way more money with AI. If I had an infinite number of employees that were experts at a task, maybe I could end up really reading every single SEC filing and finding a bunch of red flags.

How AI is changing junior roles

George Sivulka [43:49] And finding a place where the markets really are inefficient. Or if I had a bunch of expert employees with an infinite amount of time, maybe my law firm could actually litigate a case better than any other law firm in the world. And that's much harder to capture. We can't defend it with an ROI calculation, but it's part of the reason people buy.

Matt Turck [44:05] Do you come across actual cases where people will say, well, we're not going to hire X many junior consultants, analysts, bankers, because now we have this tool or similar kind of AI technology?

George Sivulka [44:30] I'll say one final story. I think we're running out of time, but I think it's a really interesting story, and it kind of talks about the future and whether or not we'll have juniors. I think we will have juniors. I've seen some people that talk about it. I haven't seen a lot of people that do it. I've seen a lot of third-party expenses, like legal fees. I'm very short Accenture and consultancies and the Big Four. But in terms of juniors, Morgan Stanley and some of the folks there always claim that they invented the analyst.

George Sivulka [44:58] The story behind that is one day they got a bunch of computers in a computer room, and they said, hey, we don't know how to use these computers—all the bankers. They hired a bunch of kids that were nerds from Columbia, and they brought them down into the computer room and they said, hey, we're going to have you use the computers. And the person that made that decision came back to Morgan Stanley, whatever, 20, 30, 40 years later, and said, hey, you still haven't figured out how to use computers?

George Sivulka [45:31] And I think the intuition or the underlying sentiment there is, when technology is created, like the computer or like AI, and it's a true revolution, you end up actually having lots of people come in and do those jobs, like do jobs related to the technology. So I firmly believe that being an investment banking junior won't look the same as it looked five years ago. I definitely believe the same for investors, for lawyers, for everyone else, but I actually don't really think that that will decrease the amount of jobs.

George Sivulka [45:48] I think that someone like Morgan Stanley will claim to invent the prompt engineer, and we can laugh about it 50 years later.

Matt Turck [45:54] Actually, a couple of last questions from me, and then I'll open it up to folks. Can I ask a personal question?

Leadership and perspective as a young founder

George Sivulka [45:55] Always.

Matt Turck [46:07] So, you are, in the grand scheme of things, incredibly young, and you're building an incredibly impressive company. Perhaps that's inspirational for anyone in the room that's thinking about building their own company. How do you do it? Maybe not in front of investors, because investors like very young founders, but you're going to go to the top people at some huge asset manager and whatever, and you're going to say, "Hey, I'm going to revolutionize your business."

Matt Turck [46:23] How do you lead as a young founder?

George Sivulka [46:46] I think naivete is a superpower, because it's incredibly hard and a terrible existence to be a founder. Everyone always says, like, "Oh, I wouldn't do it if I knew what it was like," and I think that's actually true. And so I think that young people end up changing the world because you just don't know how hard it will be. And then you're like, "Oh, I'm here. I don't have any other option." But I think it definitely demands a lot.

George Sivulka [47:15] I'm not the person that does everything at Hebbia. I do a lot of things at Hebbia, but we have a lot of talented folks later in their careers that I learn from every day, that I'm incredibly fortunate to work with, and that I'm incredibly grateful believe in me. But I'm still learning about sales playbooks and information retrieval and all the rest from people that I love working with. So I think that's the superpower, is the team, always.

Hebbia’s roadmap: Success in the next 3 years

Matt Turck [47:25] Great. And maybe just to zoom out to finish, next two to three years, whether that's product roadmap or what have you, what does success look like in three years from now?

George Sivulka [47:54] I think the only thing that I care about is, again, putting capable, truly capable AI in the hands of as many people as possible. I don't think that's a chatbot. And in three years, if the same amount of people that use chatbots today are using much more adept agentic applications and driving them to do things, to create value, to start businesses, to discover new information, that's what gets me going. That's my dream. I really want to at least play a small part in that.

Matt Turck [48:01] All right. On this note, George, this was fantastic. Thank you so much.

George Sivulka [48:03] Thank you guys for coming. Appreciate it.

Matt Turck [48:24] Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.