Making AI Work: Fine-Tuning, Inference, Memory | Sharon Zhou, CEO, Lamini

The MAD Podcast with Matt Turck · with Sharon Zhou, Co-founder and CEO, Lamini

Sharon Zhou is the Co-founder and CEO at Lamini. We cover why general-purpose models are pretty good at everything but perfect at nothing for proprietary facts, how memory tuning raised a Fortune 100 text-to-SQL system from 50% to 95% accuracy in days, and why agents making 30 LLM calls cannot realistically deliver real-time latency.

Watch on YouTube

Chapters

  1. 2:18 — The state of the AI market in July, 2024
  2. 10:51 — What is Lamini?
  3. 11:43 — What is Inference?
  4. 15:36 — GPU shortage in the enterprise
  5. 18:06 — AMD vs Nvidia
  6. 22:10 — What is Lamini's final product?
  7. 25:30 — What is Memory Tuning?
  8. 29:01 — What is LoRA?
  9. 32:39 — More on Memory Tuning
  10. 35:51 — Sharon's perspective on AI agents
  11. 40:01 — What is next for Lamini?
  12. 41:54 — Reasoning vs pure compute in AI

Transcript

The state of the AI market in July, 2024

Matt Turck [1:32] Sharon, welcome back to The MAD Podcast. You and I did a great episode just a few months ago, I guess last year.

Matt Turck [1:59] But in the AI world, a few months is like seven years on Earth or something. So let's do it again this time. We're going to talk about Lamini, we're going to talk about the markets, but let's stop here and there to talk about specific concepts that may be interesting to a broad audience that's interested in AI but may not be super into the details, into the technical stuff. So let's do that like we did last time.

Sharon Zhou [2:00] Awesome.

Matt Turck [2:00] Sounds good. All right.

Sharon Zhou [2:02] Thank you so much for having me back.

Matt Turck [2:34] Yeah, absolutely. So maybe let's start at the high level. I'm curious, from the perspective of a founder who's very much in the proverbial trenches of AI, specifically enterprise AI, day in and day out, what's the current mood in July 2024 as we record this? It sort of feels like the AI hype is slowing down a little bit. What are you seeing in the markets? How are customers responding? What are they doing? What are you seeing?

Sharon Zhou [3:00] It turns out having a non-negative margin is an interesting thing, isn't it? Who knew? Okay, so a few things. I think one is from the customer perspective, from enterprises. We sell to enterprises. I think what's going on there has been extremely exciting, but enterprises are actually starting to structure their organizations, having centers of excellence who are tackling this problem, being able to onboard different products internally. And I would say the sentiment is, hey, we've tackled the shallow use cases.

Sharon Zhou [3:25] We've been able to put some of those in production, but now we're thinking about what's next. Is this technology really only going to help me compose email? Or is it going to do something a little bit more? Is it going to help me leapfrog my industry, or is it something in between that? What's the next step? What are the steps I need to take to get there? And so I'm starting to see deeper use cases.

Sharon Zhou [3:35] I'm starting to see people tackle those deeper use cases across both their data science and engineering and even infrastructure teams. Yeah.

Matt Turck [3:41] What's an example of a deeper use case? So those shallow use cases, that's like email search, I guess, that kind of stuff?

Sharon Zhou [4:03] Yes, yes, yes. Things that for every model ping, it might not be necessarily immediate ROI. So I was talking to a large insurance company this morning, and they said, well, we have all of these essentially field reps, and they have extremely high turnover there, and we want to be able to train them up, right? We want to be able to train them up very effectively, and it needs to be grounded in these real facts about our business.

Sharon Zhou [4:23] And right now, we have these interesting demos. We have it kind of in demo land, but it's not in a place where we can realistically deploy them.

Matt Turck [4:24] Mm-hmm.

Sharon Zhou [4:45] We're not in a place where it's actually affecting thousands or tens of thousands or even hundreds of thousands of people. And that's like one step deeper. And then another step deeper is thinking about, okay, well, let's just think about drug discovery, right? Drug discovery takes a decade with a 90% failure rate today. But imagine if a company could pull that in and do, I don't know, like only 30% failure rate in one year.

Sharon Zhou [4:55] Like, that would be a dramatic improvement.

Matt Turck [4:55] Yep.

Sharon Zhou [5:12] And that would be very, very interesting. That would be a true leapfrog where this company would completely shatter that field, right? Like, because you would succeed at a completely different timescale and at a completely different failure rate. So I think that future is coming, but it's not necessarily coming as fast as ChatGPT entered everyone's phones and made it possible for us to send a text message and be able to get some form of an intelligent answer back.

Matt Turck [5:42] Yeah. So what are big enterprises doing currently? Are we still in a stage where they brought in the consultants, they built the center of excellence? I mean, what are they doing, and what would you recommend they do to move as fast as they need to on generative AI?

Sharon Zhou [5:49] So one thing that they've already done, I see, is I think compute infrastructure has been laid out more in place. I think budgets are there.

Matt Turck [5:50] Mm-hmm.

Sharon Zhou [6:17] I think the next thing around organizational structure that's extremely important that I think people don't fully realize is relating the technical piece of building out a successful model with the development piece of this model. And what I mean by that is generative AI is famously very hard to evaluate. We have no idea what's good, better, best, unless someone who's an expert in understanding that use case can tell you that, right? Like, I can't tell you a medical generative AI model is actually that good because I'm not a doctor, right?

Sharon Zhou [6:51] I can tell you when it's at, like, toddler-teenager stage, but by the time it's getting an MD, I'm out. I'm not there. So being able to have those people sit closely with the development teams actually accelerates those use cases much more quickly. And that becomes an organizational problem. That becomes a very, very big difference I see across some enterprises where those teams are closer together, so those use cases can get out much more quickly, and then other enterprises where those are much more disjoint today.

Sharon Zhou [7:15] So they need to reorg to be able to actually get those closer together in order to deliver those applications. And I think that's why realistically we've seen many more code agents and applications around code, maybe text-to-SQL. We've seen that a lot because the developer can also evaluate the outputs, right? They are one and the same person.

Matt Turck [7:16] That's true.

Sharon Zhou [7:30] And so that makes development much faster because, again, that is a famously difficult problem in generative AI from back in the research days. That was my PhD dissertation, actually, to today for realistic applications.

Matt Turck [7:54] Yeah, very interesting. And what are you seeing, again, for enterprise customers in terms of, like, oh, let's just bring in OpenAI on top of Azure or whatever to do these use cases, versus let's play with open source, our own models? And how's the thinking evolving?

Sharon Zhou [7:58] Yeah, so I think it's pretty clear that it's going to be a hybrid world, right?

Matt Turck [7:58] Mm-hmm.

Sharon Zhou [8:26] And I think the general-purpose models of the world are very good at these quote-unquote shallower use cases that aren't necessarily business-specific, right? Like, they're not necessarily going to be your proprietary model or heavy proprietary differentiation using AI. But they are going to be important to composing email, maybe writing certain types of basic code that aren't relevant to deep codebases that are very custom. So I think that will always be there, and there will always be a best general-purpose model, whether it be OpenAI's or Anthropic's or Google's, et cetera.

Sharon Zhou [9:04] And then I'm starting to definitely see a shift of, okay, well, this is great. This is an absolutely amazing technology. We've gotten to develop advanced RAG and prompt engineering solutions on top of it. But it's still not enough. These models, when they're general, they're optimizing for what's known as generalization error, or the average error across all examples it sees on the internet. And as a result, it's pretty good at everything, but it's perfect at nothing. And despite our flaws, Matt, you and I, we're actually perfect at some things.

Matt Turck [9:10] Are we?

Sharon Zhou [9:39] Yes, I remember your name. You remember mine, presumably, and your own birthday, maybe, and maybe revenue number last quarter. I remember the liability number in my insurance claim, for example, if I were an insurance company. So I want the LLM to remember these facts with crisp accuracy, right? With factual accuracy where there's no alternative at all. And that means being perfect at some things.

Matt Turck [9:39] Yeah.

Sharon Zhou [9:55] That means being perfect at recalling whether it be a column in my SQL table or, again, that revenue number in my earnings report. Like, that is incredibly important for the reliability of these models. And so slightly right is not the same as right for these facts.

Matt Turck [9:56] Yep.

Sharon Zhou [10:23] And that is where I think the next frontier is of, okay, well, now we've been able, with memory tuning, which is what I've been working on, to remove those hallucinations, to remove that, and actually get these models from not necessarily being general for everything and instead of being pretty good at everything but perfect at nothing, to be actually perfect or near-perfect at some things and still pretty good at everything else. And I think that's where we're going.

Sharon Zhou [10:29] And I think that's extremely relevant to the enterprise.

What is Lamini?

Matt Turck [10:56] Yeah, great. All right, so memory tuning is definitely something I want to spend a good amount of time on in the discussion. Maybe for anybody listening to this that didn't listen to the prior episode, which they really should do—yeah, why not?—where we describe your very fascinating personal background and all the things, maybe a quick reminder on what Lamini does, like the 30-second pitch.

Sharon Zhou [11:20] Okay. Yeah, 30-second pitch: we're an integrated inference and fine-tuning platform for enterprises to be able to run factual LLMs. So essentially, LLMs that don't hallucinate on their proprietary data within their secure walls. So we can deploy on-premise, air-gapped, no-internet sites, so in the most extreme cases with extremely high security.

What is Inference?

Matt Turck [11:54] Okay, awesome. Very good. So since you and I last caught up, one big evolution is that you are now offering inference. So when we first talked, you were very focused on fine-tuning, which is part of the training world. And you now added the other side, which is inference. And maybe, as I said upfront, we'll define a bunch of things. So for anyone that's not super in the weeds of AI, what is inference?

Sharon Zhou [12:24] Inference is when you run the model. When you send ChatGPT a text, you're running inference. So oftentimes, we're running inference of the model. And yeah, we offer inference now. It had always been bundled. I almost view the future of a world where we will get to continual fine-tuning, continual inference, and continual requiring that integration, that bridge between the two, because they will both be happening at very fast speeds.

Matt Turck [12:24] Mm-hmm.

Sharon Zhou [12:45] But essentially, my view of inference today, from a market perspective and startup market perspective, is that it's a race to the bottom today for cost. And I don't think that's a controversial statement at all. I think people know that it's getting priced lower and lower. So we just offered 40 million tokens for free. And I think, in a sense, we can, as a startup, be able to say, hey, actually, our differentiation is in tuning the model, so we can actually just offer inference completely free.

Matt Turck [12:57] Mm-hmm.

Sharon Zhou [13:14] And so this is essentially a way for us to initially get customers and initially get them started, and then future-proof them as well, enable them the flexibility to be able to then use these models, but also improve them over time for those future use cases if they are only on inference today.

Matt Turck [13:25] Could you map the inference market for us? Like, who does what? There seems to be a number of different companies doing different things. I'm thinking of Modal, this open-source project like vLLM. Who does what?

Sharon Zhou [13:46] Yeah, I'll map that out and kind of focus on what are the key differentiators. So one key differentiator that I think we often see on benchmarks is around latency. So latency is how fast the model responds to a single query. That makes it so things are real time when you chat with ChatGPT, and that can move faster. Unfortunately, I think what people don't realize is that the way to get better latency, like significantly better latency, is actually in the hardware.

Sharon Zhou [14:22] And that's why we see Groq, G-R-O-Q, be able to exceed all these GPU-based inference platforms significantly, by a significant margin. And I think I saw in a VentureBeat article about 280,000 users now using that. So props to them for getting that there. So I think for latency, if you want to compete with speed of the LLM's response to you for a single response, then you need to compete at a hardware level. The next thing is around security.

Sharon Zhou [14:52] I think security is interesting because are you running it on a cloud that someone else is managing for you, or are you running it on your own GPUs, whether that be your own GPUs in a VPC or your own GPUs on-premise? So I think that there's a distinction there of what level of comfort do you have in sending data over to another party? And I think I see AI natives and startups largely using the hosted solutions. So managed solutions, we have one as well, kind of on the cloud, to make it very easy for people to test out.

Sharon Zhou [15:00] And then also—

Matt Turck [15:01] And that's the Togethers and the Modals.

Sharon Zhou [15:24] The Togethers and the Modals of the world. And then on the other side of things is being able to deploy that stack on-premise or in any hybrid environment and be able to fully control that and have visibility through that. And so we also do that. That's our main way of deploying things. But essentially, you can think of that as, I think NVIDIA's Triton Inference Server also does that. And so I think these are all powerful ways of getting high performance.

Sharon Zhou [15:33] vLLM is an open-source project that does that for high performance as well.

GPU shortage in the enterprise

Matt Turck [16:01] Yeah. How does that work for enterprise customers? So in a world where you have a massive GPU shortage, do they need to procure their own GPUs and then run an inference platform like Lamini and bring their own models? I guess that's what you do for them. But are they able to get the GPUs they need?

Sharon Zhou [16:28] How does that work today? Actually, I'm seeing the GPU shortage go away at the company level, meaning companies are able to procure enough compute. Enough is a strong word, but they're able to procure compute at some level to work with, to fine-tune and run heavy inference jobs for these models. However, I think individual teams, that varies, right? Being able to get it from one business unit, get it from central IT, I think there's still a process there. It's kind of like getting data access.

Sharon Zhou [16:36] Compute access is also gated in a way that is variable across companies, I would say.

Matt Turck [16:49] Okay, but if you're a large company—just to play it back—if you're a large bank or Pfizer or whatever today, July 12th, 2024, when we're recording this, you're actually able to get GPUs?

Sharon Zhou [16:49] Yes.

Matt Turck [17:06] Because we're in a world where there are announcements about VC firms, which I think is a very smart strategy, stockpiling GPUs and all the things. But what you're seeing in the market is actually, if you're a large enough company, you do have access to GPUs compared to last year.

Sharon Zhou [17:32] Last year at this time was absolutely insane. That's why we threw up our own cloud, because large companies with multibillion revenue numbers could not get a node from AWS despite their accounts being tens of millions or hundreds of millions with AWS. And they said, "Oh, I got one node for one month. That's what my rep said." And I thought, okay, well, that's not great if that's C-suite asking the rep. There's probably something wrong here.

Sharon Zhou [17:44] So that's no longer the case. I think people have far more compute. Still probably not enough compute to satisfy demand, but it's not totally blocked.

Matt Turck [17:44] Yeah.

Sharon Zhou [18:05] It's not in that ridiculous state anymore. And I think there's a graph of this showing that A100s in particular are pretty available. Obviously, the AMD chips that we also agnostically work with, the MI300 and MI250s, those are available. H100 is still kind of a little bit harder to get, but you can get started very easily with any of those other chips.

AMD vs Nvidia

Matt Turck [18:22] Okay. And you mentioned the GPU-agnostic part of the business. So you've been working with different types of GPUs, but last year, part of your positioning was to also work with AMD. So what's the update?

Sharon Zhou [18:25] Yeah, well, today we did announce something with NVIDIA.

Matt Turck [18:26] Just today?

Sharon Zhou [18:28] Yeah, just a few minutes ago.

Matt Turck [18:30] It's like The MAD Podcast breaking news.

Sharon Zhou [18:32] Exactly. Exactly.

Matt Turck [18:33] We're now a news organization.

Sharon Zhou [18:37] Well, we put that out with NVIDIA, which is awesome.

Matt Turck [18:37] Yeah.

Sharon Zhou [19:00] And we're really here to make it easy for developers and people building on this technology to have a way to not have to edit their code at all and be able to run agnostically and scale agnostically on multiple types of compute. I think that's going to be extremely valuable moving forward. And I see that, for advanced organizations, I see them doing it today. I think for many organizations, that will be an emergent value prop for them as they scale.

Matt Turck [19:30] Okay. All right, so let's say I'm a large customer. I'm a Lamini customer. I bring my own GPU, I get access, and then you enable me to then bring my own model, or you have open-source models built into the platform. Let's say I want to use Mistral. Is that built into Lamini?

Sharon Zhou [19:54] Yeah, we have Hugging Face's model loader integrated. So any LLM, open-source LLM there, can be used, downloaded. If you're online, you download it automatically from Hugging Face. If it's an offline kind of deployment with no internet, we download those beforehand, or there's a shared file system you can put those models into. We have a few companies who are pre-training their own models and using those and then fine-tuning. But yeah, for the most part, fine-tuning is taking a pre-trained model, a foundation model, and adjusting its weights further.

Sharon Zhou [20:13] But taking something off the shelf is a very reasonable place to start because Zuck decided to put in billions of dollars upfront for us, and we can reap the benefits of those savings.

Matt Turck [20:14] Thanks, Zuck.

Sharon Zhou [20:24] Yeah, thank you. Thank you for the potential scorched-earth method of LLM strategy. Keep at it. I do hope he keeps at it. Yeah.

Matt Turck [20:33] While we're at it, talking about definitions, you mentioned pre-training. What is pre-training versus training?

Sharon Zhou [20:57] So pre-training is a form of training, actually. And I think people often think of those similarly today for LLMs. But essentially, it's getting the LLM to read the entire internet one word at a time, or one token at a time. And it's just autocompleting the internet, right? And so that's what it's tasked to do. It doesn't really know how to answer questions. So if you ask it, "What's the capital of France?" it's going to respond, "What's the capital of Spain?"

Sharon Zhou [21:24] Because it thinks it's in maybe a survey context. And so after pre-training, most foundation model companies, for example, like Meta with Llama and OpenAI and Anthropic, will do something called instruction fine-tuning, which teaches the model how to follow instructions so that when you ask, "What's the capital of France?" it will say Paris. And so getting it there so that it's chatting with us. And then the next step we often take customers through is memory tuning, because that's not a general skill for the model; it's more specific to your data.

Sharon Zhou [21:58] So, to be able to embed facts of your data into the model, to memory-tune the model so that it can recall those facts almost deterministically within its probabilistic context. So that's kind of the last step there. And then, of course, with inference, you can apply RAG and prompt engineering on top of that model in any of those stages to be able to adjust it in real time to output something that you need. So that's kind of the layout of how these models are tuned and trained and adapted and turned into these magical systems that we use today.

What is Lamini's final product?

Matt Turck [22:25] Okay, fantastic. And so, still talking about Lamini in the enterprise, we got the GPUs, we got the models, we touched upon, I guess, the data a little bit through RAG. The output is what? It's a JSON that people can bring into their applications. What comes out of Lamini?

Sharon Zhou [22:28] Oh, what comes out of Lamini? So, we are a stack.

Matt Turck [22:29] I mean, magic, of course.

Sharon Zhou [22:40] No, not magic. Actually, I want to make sure it doesn't feel like—it can feel magical, but hopefully developers can dive in and actually understand mechanistically what's going on.

Matt Turck [22:40] Absolutely.

Sharon Zhou [23:03] So our customers are able to fully own their models, right? We don't own the models at all. We're just the engine that runs it. Similar to—you can think of it as a Databricks or Snowflake model of compute, like being able to manage that compute. We're a stack on top that people are able to install on top of Docker, Kubernetes, or even bare-metal compute. And developers are able to interface with it through familiar REST APIs or an optional Python client, et cetera.

Sharon Zhou [23:19] And so, be able to actually work with these models locally or in whatever environment they're in, be able to adjust them, and be able to then run them at scale.

Matt Turck [23:32] And maybe to close on inference, I was reading—or I'm reading right now—that Lamini delivers 52 times more queries per second than vLLM, which we touched upon. What's the secret sauce?

Sharon Zhou [24:00] Yes. So a huge part of this is around being able to fully utilize GPUs. I think what we found was a lot of the open-source packages out there, while incredible—and I root for those teams every day—they're not necessarily optimized in the multi-GPU case, multi-GPU scenario. So if you're using more than two GPUs, which nearly everyone I know is, then fully utilizing the GPU is actually incredibly important to eke out every bit of it you can. This is a very expensive piece of hardware you just bought.

Sharon Zhou [24:32] If you're only using a small percentage of it and you're running open source, it actually is more cost-effective to run something—to pay us to install the inference, right? And to run that at high performance so that it is fully utilizing your GPUs at scale for more than two GPUs. So I actually find that dichotomy pretty interesting. So throughput is—you can think of it as batch latency. It's essentially how many queries you can put in. It's slightly separate from latency itself of a single query.

Sharon Zhou [24:55] But that, again, is, I think, a way companies are differentiating today as well. And on inference, actually, I think the final thing that people are highly differentiated on is close to something with function calling. So being able to output structure is really important to programmatically use these LLMs. So developers really love it when you can output something that is 100% schema accuracy, 100% format faithfulness, essentially, from the model—be able to output a list of ints where all the ints are my product IDs or are my user IDs, right?

Sharon Zhou [25:29] So being able to make sure that it can actually output that right format so that I can then use it downstream for something else, especially in an agent case, for example. So that's something else we offer through our inference service to actually make it 100% by re-engineering the decoder of any LLM.

What is Memory Tuning?

Matt Turck [25:45] Great. Memory tuning. Yes, we alluded to it. It feels like you all had a major breakthrough in terms of research. What is it? Why does anyone need it? And how does it work?

Sharon Zhou [26:10] Yes. So maybe stepping back, these models are pretty good at everything, perfect at nothing. And it's because they're reading the whole internet, they're trying to just reduce average error, average error over everything. But average error means you're not actually perfect at those facts. And it turns out we actually want something closer to deterministic on those facts. And so how do we mix what is probabilistic and really good at understanding similarities, where we want hi and hello to mean kind of the same thing, but we also want those facts to be absolutely perfect, where if you get your birthday off by a day or two, that's actually a major issue.

Sharon Zhou [26:53] If you get your revenue number off by a zero, that's a major issue. But from the model's perspective, it's not. So how do you make sure that, for those facts, those specific facts, your business or otherwise, there's no alternative for the model? And so memory tuning does that. It essentially, we say, brings the loss to zero. So it brings the loss to zero on all of these facts so that there is no alternative. It cannot even pick something that is similar.

Sharon Zhou [27:21] So something slightly correct is incorrect, as opposed to being, oh, kind of similar, let me maybe consider that. So memory tuning is a technique to do that. It's very computationally expensive to run memory tuning. So we've essentially optimized it so that it is computationally feasible by taking any kind of open model and tuning the adapters. So tuning something like a LoRA, QLoRA, et cetera, on top of the model, and tuning it in such a way that it's a mixture of experts.

Sharon Zhou [27:50] It's a mixture of expert adapters. I think there's been a trend of mixture of experts. GPT-4, Mixtral are all a mixture of experts. Those are eight experts so that you can train a bigger model, but then when you run it, it only uses part of the model, one of the experts, one-eighth of it. What memory tuning is doing is taking that to the extreme and doing it on the adapter side of things, so doing it on top of the model with LoRAs.

Sharon Zhou [28:14] And we tune these memory experts, and we have a million memory experts or more. And so the idea is instead of eight, now you have a million, and you essentially get a sparsely activated, heavily sparsely activated model. So you can scale the model to be incredibly large. I even think there's a future where these models can be 100 billion parameters, but have that intelligence of 100 billion parameters, but then have the speed, latency, and cost of something that's still 1 billion or 7 billion parameters.

Sharon Zhou [28:49] So very, very fast because it's only using part of its brain when it's responding to you. Kind of like how we operate. So I think that's a trend, and we just took that to the extreme, and it's highly effective for bringing that loss to zero in a reasonable amount of time, such as a couple of hours.

Matt Turck [28:57] Okay, fantastic. So you cannot resist again to make this interesting to a broad group of people.

Sharon Zhou [28:58] Yeah, please.

What is LoRA?

Matt Turck [29:10] With your academic hat on, let's do 45-second definitions of some of those terms. So let's say, what is LoRA?

Sharon Zhou [29:40] Oh, LoRA. Okay, so when you tune a model, you have this giant model, right? And if you want to tune all of the weights of the model, which is traditional fine-tuning, it's very expensive, very heavy to change, to do all those multiplications, additions, et cetera. So that's very, very expensive. LoRA is an efficiency technique to make fine-tuning extremely efficient. It's part of the classic category of parameter-efficient fine-tuning, PEFT. And instead of training all those weights, all those parameters, you can tune only an external set of weights instead.

Sharon Zhou [30:16] And sorry, not a subset, an external set of weights instead. And then at inference time, fuse those weights using, I'll just say, math, back into the model so that it is the same latency. So it's not just external, like extra weights, and the model has to compute all of that when you ask it a question. You fuse it back into the model, so it is the same latency when it comes back to you with a response.

Matt Turck [30:16] Right.

Sharon Zhou [30:43] So it's an incredible technique. I think for something like a GPT-3-level model, it's a 10,000x speedup in efficiency while losing nearly nothing at all in accuracy. And for enterprise use cases, that's really not the real bottleneck in accuracy. So we see it as just a performance improvement. And I believe folks are all trending towards fine-tuning with LoRA, since it's kind of what makes sense.

Matt Turck [30:47] And in the same vein, what is mixture of experts?

Sharon Zhou [31:06] Oh, yes, mixture of experts. So mixture of experts is essentially—well, let's step back. Before mixture of experts, we have one giant model. Okay, so it's one giant model. And when I ask it, "Hi, how are you?" every single weight of the model, the entire brain of this model, has to think, think, think and respond, "I'm good, thanks." And then if I give it a really, really, really hard math question or really, really hard patient diagnosis, it will also have to use all of its brain and then give me the diagnosis.

Sharon Zhou [31:39] Right. And so that doesn't seem to make any sense to do that. That seems like overkill. So mixture of experts is getting at the sense of, okay, well, how about we have maybe eight experts instead, and they all kind of come together. They all still train together, and there's a router that routes between them so that when I ask, "Hi, how are you?" maybe the expert who's just good at regular conversational English prose will say, "I'm good, thanks."

Sharon Zhou [31:59] And another one who's good at medicine will give me that diagnosis from the patient. And the benefit of this is that these experts are much smaller than the whole thing. It's only using part of the AI brain, right?

Matt Turck [31:59] Mm-hmm.

Sharon Zhou [32:24] So that makes it so it's much faster when it responds. It makes it much cheaper when we run fewer computations on the GPU. And it's also much less memory-intensive because it's only running part of that model there. So, yeah, that's one interesting thing about mixture of experts. And of course, I kind of like seeing this taken to the extreme with LoRAs together. And I do think that's the future, so we can get something that is incredibly smart, incredibly huge, but with the latency, costs, and speed of something tiny.

More on Memory Tuning

Sharon Zhou [32:39] No more big model versus small model paradigm. It's potentially one and the same.

Matt Turck [32:50] So back to memory tuning. Help us understand: is that something that's in production today, the product, or is that research?

Sharon Zhou [32:51] Yes.

Matt Turck [32:57] And if it's in production, how is it working so far?

Sharon Zhou [33:28] Yes. So it is in production today for enterprise customers. We put out a case study with a Fortune 100 company, and they were doing text-to-SQL. They had spent 6 to 9 months with 3 data science teams working with advanced RAG and instruction fine-tuning, as well as prompt engineering, and could get to about 50% accuracy. They had very complex schemas, and with memory tuning, in just a few days, they were able to reach 95% accuracy on their same evaluation set. So, very big change, because the exact problem they had was that the model was hallucinating on, for example, month ID 2024-04.

Sharon Zhou [34:05] Turns out the model, you could not get it to say anything but April, but for their database, it was January, right? Maybe there's a fiscal year thing going on, right? So it's very custom to that company, and they could not, with as much RAG, as much, "please, 04 means January," get the model to do that. So this is actually enforcing that inside of the model. And it's very exciting. And I'm seeing basically a new frontier of accuracy emerge for these models and for use cases for enterprises.

Sharon Zhou [34:29] And so that'll be really exciting. I think there will be completely new use cases emerge from this that are now possible, that were impossible before, and it's at a lower cost and latency. So that'll also help with getting real ROI on these use cases.

Matt Turck [35:00] Yeah, that feels like an absolutely huge breakthrough. Hallucination has been the main issue. And if you could get to that level of accuracy and make things like text-to-SQL actually work, that is incredible. Congratulations on that. Okay. And the trade-off, you said, is you've made it computationally feasible, but it's still computationally intensive.

Sharon Zhou [35:25] So I would say today memory tuning is probably easier than advanced RAG, but still harder than simple RAG. And from a speed perspective today, it's still a couple hours, like 2 hours, to be able to run a memory tuning run, whereas building a RAG index is on the order of minutes. So we're working on bringing that number down to minutes so that it feels like building an index. And memory tuning, to be clear, is actually super inspired by not necessarily RAG, but the same parent area called information retrieval.

Sharon's perspective on AI agents

Sharon Zhou [35:51] And it's essentially putting that information retrieval piece inside of the weights of the model, in a way that is extremely efficient for that AI brain to retrieve the right memory experts for the right facts that they're trying to respond with.

Matt Turck [35:58] Okay, so the other big theme that everybody's talking about in the enterprise these days is agents.

Sharon Zhou [35:59] Yes.

Matt Turck [36:06] Where does that fit in the Lamini vision, and what are your general thoughts on the world of agents?

Sharon Zhou [36:31] Yeah, I think it's a spicy topic. So, we support a lot of agent use cases today, and obviously the LLM is behind running those agents. The case studies on a SQL agent, right, able to ping a database with the SQL it's generated and be able to respond back to the user. My view of agents is that it's kind of object-oriented programming where it's taking a different center. Instead of looking at each LLM, each LLM call, it's saying, okay, here's something that is almost like something I understand, like an object or a person, and being able to say, okay, it can do all of these tasks.

Sharon Zhou [37:07] And all of those tasks are LLM calls. So I kind of think of an agent as something that does more than one LLM call and is thinking about it from this center of gravity of, how person-like can I get this so that I understand it in my worldview? It's interesting because historically, AI researchers, we haven't really thought about it from the lens of agents. I think that really has emerged from the community of developers, and because I think code agents became incredibly interesting and hyped up.

Sharon Zhou [37:37] A little bit. Just a little bit. And I think from a marketing perspective, and an explainability perspective, I think it's actually incredible because it explains how people view the world and what AI could do. Now, there is a gap between what it can do and those goals and potential fantasies. So I think bridging that gap is probably an important piece of it, and understanding that, at the end of the day, the most computationally expensive and the most important piece of that agent is the LLM.

Sharon Zhou [38:20] Like, that is the hardest piece of it. And I am a little bit worried about those who call an LLM like 30 times within a workflow and still expect real-time latency on them because, realistically, that's probably not the approach you want to take if you expect your agent to be real-time, like a person, and then you do 30 LLM calls, which all require extremely high intelligence to be able to understand what's going on. And so I think there will need to be specialization.

Sharon Zhou [38:33] And we've started seeing people both memory-tuning and fine-tuning these models so that, instead of 30 calls, three calls or even one call make it much more real-time.

Matt Turck [38:48] And is that a question of it's just too computationally intensive and therefore expensive and then slow? Or is it a question of the error, the potential hallucination, compounds? Or is it both?

Sharon Zhou [39:13] It's both of those things. And I think people are addressing error today by adding more calls to the model, filtering out the requests. And I don't think that'll work for serious production use cases. It makes it harder to work with that entire blob. But for a prototype, I think it's perfect. I think it's great. Yeah, show that it is doable or almost doable, right? I think that actually is able to help you get the right data to then maybe even fine-tune or adjust the model to do it in one shot, in one call.

Matt Turck [39:36] Okay. All right, all right. Well, it's been wonderful. Maybe to close, what's next at Lamini the next few months or year, until you come back to The MAD Podcast?

Sharon Zhou [39:41] Until I come back? Oh, what are my predictions of what I'll say next time? Yeah, no, no.

Matt Turck [39:52] Well, hopefully we can do this on a regular basis. I very much enjoy those conversations. So, joke aside, what is next for Lamini?

Sharon Zhou [39:54] That was a joke.

What is next for Lamini?

Matt Turck [40:02] All right, well, it's on record now. You have a permanent invitation to come back. So what is next for Lamini?

Sharon Zhou [40:30] So what is next? I think we want to drop the speed of doing memory tuning significantly because we've seen huge value in the market for this technology, obviously with reducing hallucinations being a major gap in production use cases. But dropping the speed, I think people sometimes think, oh, performance, whatever. Serious drops in speed and in performance result in the ability for completely new interfaces and tasks. So if instead of hours it takes minutes, if it takes seconds, I believe in a future where we're continuously fine-tuning these models, where it's as easy as prompt engineering, and these models continually improve.

Sharon Zhou [41:15] It actually baffles me that when we do text ChatGPT and we give it all this feedback in my conversations with it, that it's not actually adjusting the weights, because it is quite difficult to do that. And so to make that significantly easier, and then also faster, those are two important pieces of bringing that future out. And I think that is the full capability of these models. Because today, when people are running RAG or prompt engineering, those are search. That's not AI.

Sharon Zhou [41:19] It's like keeping the AI frozen and fixed.

Matt Turck [41:19] Mm-hmm.

Sharon Zhou [41:43] And so I'm curious about a world where the AI is able to continually evolve with us, just like we learn. It's not like our brains are not plastic, right? So a brain that's plastic is going to learn much, much more versus just be kind of adjusted at a local scale. So I'm excited about that. And I think, yeah, I'm excited to see that actually come to light. And I see people starting to do continual fine-tuning in a way that is maintainable.

Sharon Zhou [41:52] So to keep models refreshed and up to date.

Reasoning vs pure compute in AI

Matt Turck [42:09] And are you, while we're on the topic, in the camp of ultimately needing to add reasoning, symbolic parts to the raw power of LLMs?

Sharon Zhou [42:11] Or is it just data and compute?

Matt Turck [42:13] Yes. Where are you on that debate?

Sharon Zhou [42:46] I think part of my heart wants it to be that way, where we understand enough that we can symbolically include something in there. But my sense today is there's nothing that's convinced me that that is good enough. Instead, the new architectures that come out that are essentially better are all about taking advantage of compute better. And that is what transformers were, and that's what some of the new models coming out and new architectures are as well. So I'm excited to see what the future bears, but I don't think it's going to be the exact architectures and how we frame it, how we understand it today.

Sharon Zhou [43:29] Yeah. One in particular that I'm interested in is diffusion models. I used to study diffusion models, which would significantly reduce latency by a very large margin. So I'm interested in seeing—I saw a masked diffusion for text generation that was very promising in the research realm. And so I'm kind of curious about that. Improvements that are not necessarily incremental. And when fields kind of collide, I think new things happen. Yeah. Great.

Matt Turck [43:32] Well, fantastic. Sharon, thank you so much for doing this.

Sharon Zhou [43:34] Thank you so much for having me again.

Matt Turck [43:55] Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.