What You MUST Know About AI Engineering in 2025 | Chip Huyen, Author of “AI Engineering”

The MAD Podcast with Matt Turck · with Chip Huyen, Author, AI Engineering

Chip Huyen is the Author, AI Engineering. We cover why evaluation is the biggest bottleneck to AI adoption, why RAG remains necessary even with million-token contexts, and why training data for agent planning is difficult because human-efficient plans differ from AI-efficient ones.

Watch on YouTube

Chapters

  1. 2:45 — What is new about AI engineering?
  2. 6:11 — The product-first approach to building AI applications
  3. 7:38 — Are AI engineering and ML engineering two separate professions?
  4. 11:00 — The Generative AI stack
  5. 13:00 — Why are language models able to scale?
  6. 14:45 — Auto-regressive vs. masked models
  7. 16:46 — Supervised vs. unsupervised vs. self-supervised
  8. 18:56 — Why does model scale matter?
  9. 20:40 — Mixture of Experts
  10. 24:20 — Pre-training vs. post-training
  11. 28:43 — Sampling
  12. 32:14 — Evaluation as a key to AI adoption
  13. 36:03 — Entropy
  14. 40:05 — Evaluating AI systems
  15. 43:21 — AI as a judge
  16. 46:49 — Why prompt engineering is underrated
  17. 49:38 — In-context learning
  18. 51:46 — Few-shot learning and zero-shot learning
  19. 52:57 — Defensive prompt engineering
  20. 55:29 — User prompt vs. system prompt
  21. 57:07 — Why RAG is here to stay
  22. 1:00:31 — Defining AI agents
  23. 1:04:04 — AI agent planning
  24. 1:08:32 — Training data as a bottleneck to agent planning

Transcript

What is new about AI engineering?

Matt Turck [1:21] Chip, welcome.

Chip Huyen [1:30] Hey, Matt. It's great seeing you again. Big fan of the work. Love the jokes on Twitter. So it's really nice catching up again after following you for so long.

Matt Turck [2:01] Appreciate it. So today we are going to talk about your brand new book, published by O'Reilly, which is just coming out, entitled AI Engineering: Building AI Applications with Foundation Models, which I must say is incredible work. So I spent a good portion of last weekend reading it, and I thought it was amazing. Absolutely a must-read for anyone that's curious about the AI field. And in particular, what I found really interesting is that there's plenty for technical folks. There's math, there's in-the-weeds kind of details, but equally, I found it very approachable for non-technical people, which is very hard to do.

Matt Turck [2:35] So again, really enjoyed it. Congrats. And to jump to the punchline, people should absolutely get the book. And what we're going to try today is give people a little bit of a flavor for what's in it. So obviously, we're not going to cover everything because it's 500 pages of goodness. But hopefully that will give people some kind of overview. Does that sound good?

Chip Huyen [2:41] Yeah, thank you so much. And everyone, listen to Matt. He knows what he's talking about. So appreciate it.

Matt Turck [3:14] All right, so let's jump into it. So at the beginning of the book, you make the point that while AI adoption seems new, it's built upon techniques that have been around for a while, like language models, some of which came in the 1950s, and then retrieval techniques. But at the same time, it feels like a new field. So what is new about AI engineering, and how is that different from more traditional machine learning and MLOps techniques?

Chip Huyen [3:37] Yeah, I think that's a great question, and I get asked that question a lot. It's like, okay, what is AI engineering? Is it another marketing term? How is it different from my traditional ML engineering? So there's a lot of overlap between these two roles. And I think at a lot of companies, even people with the same title can have very different functionalities. So I think any definition is a little bit fuzzy and really depends on where you work and what you're working on.

Chip Huyen [4:12] But in general, I think of machine learning engineering as when you have to build the models yourself. Before the availability of large language models or foundation models that anyone can access, if you wanted to build ML applications, you would need to build the models yourself, and only a few organizations could do that. But nowadays, anyone who wants to leverage AI to build applications can just leverage one of those amazing available models to do so. It just makes it so much more accessible.

Chip Huyen [4:41] Another thing is that before, I had thought that a small improvement of AI capabilities could lead to a small increase in the number of available applications. We have known for a long time that if we put more data and more compute, we'd get better models, right? But still, when ChatGPT came out, we were shocked. At least I was in a group chat with a bunch of my friends, and we were really shocked. The reason is that we were shocked that just a small improvement in capabilities can lead to so many applications.

Chip Huyen [5:15] At the same time, we have so many new ideas, and it's so easy for people to build applications. Just like the energy, the community is growing exponentially. It's a really, really exciting time. There are a lot of new things with that. One is that evaluation has become so much harder. Before, with a lot of traditional ML, we have, okay, if we do spam detection, we know that the output should be spam or not spam. If the model's output is not spam and the real email is spam, then we know that the prediction is incorrect.

Chip Huyen [5:50] But now, if you ask the models, you say, hey, summarize a book, and the summary looks quite reasonable, coherent, you don't know if it's a good summary or not. You might actually have to read the book to find out yourself. And also, the more intelligent AI becomes, the harder it is to evaluate it. So, for example, for math problems, I think that most of us can tell if the solution to a first-grade math question is wrong or not. At least I hope that most of us, with a lot of complaining about education going on, I'm not sure if that's still the case.

The product-first approach to building AI applications

Chip Huyen [6:21] But for the PhD-level math questions, I think very few of us can actually tell. So it's just like now AI can answer questions that's not very correct. A lot of the failures are silent, so it's really hard to evaluate. And another aspect is a question of product. I think before, you knew what you were going to build because you're part of the organization. So once you leverage data and ML, you speed up whatsoever, make more money.

Chip Huyen [6:50] So you knew what you were going to build. But nowadays, it's like you can build anything. So usually you start with a product. You start with an idea, a demo. And then if it goes really well, then you start investing in data to make it better. And then if it's cooked really well and say, okay, now everyone, we are paying too much money for OpenAI and Anthropic or Google. So now we need to build our own model.

Chip Huyen [6:58] And so now we invest into the model. So the process now is like product to data.

Matt Turck [7:08] Is that the reverse process compared to traditional machine learning, where you would start with the data and then build the model and then build the product?

Chip Huyen [7:34] Yes. So I think the process is reversed and also leads to people a lot closer, like product people and ML people and data people, a lot closer. Like in some teams, you might have, like, the same persons. So it requires engineers to have a much, much better product sense. So a complaint I often hear is just like, okay, the technical aspect of this application is really easy. But understanding what users want is really, really hard.

Are AI engineering and ML engineering two separate professions?

Matt Turck [7:51] So does it mean that a traditional machine learning engineer and an AI engineer may be two different people, two different professions? What is the overlap between both?

Chip Huyen [8:18] So I see it's very different at different companies. For example, I see a lot of companies that already have an ML team. So when they started adopting AI foundation models, generative AI, then they tasked this team to, like, okay, go and explore and build. So I think there's overlap, but I also see a lot of teams hiring separately for generative AI. So it really depends from organization to organization. One thing I do notice is that it's not a question of, like, do I use a classifier or generative AI?

Chip Huyen [8:51] It's more of an and. The vast majority of generative AI systems I've seen have traditional or analytical ML components with generative AI. Let's say, for example, a customer support chatbot. You have a request from customers, and then you want to respond to that. Maybe you employ a bunch of models. Maybe for a really difficult request, you want to send it to the strongest and most expensive model. But the easier requests, send them to maybe a locally hosted open-source model.

Chip Huyen [9:26] But some requests are sensitive, like a billing complaint, so you might want to route to a human operator, right? So when you get a request, you need an intent classifier to say, hey, what is this request about? Then you can route that accordingly. That intent classification model can be a traditional ML classifier. Another thing is that when the model, maybe a fancy model, generates a response, what if it contains some personal PII information? You might have a detector like, hey, does this contain PII?

Chip Huyen [9:53] Or not. So that can also be, like, a traditional machine learning model. So I see that used together in a lot of cases. So I do think that there's a lot of overlap. And I do think, just like people can, if you're from a traditional ML engineering background, you can also just learn more about foundation models and how to work with foundation models to become AI engineers. I also see people coming into AI engineering with absolutely no ML background because you can just do an API call, and a lot of people can't really explain gradient descent.

Chip Huyen [10:20] And not necessarily. I think it's a big debate on whether people need to know that to become a good AI engineer or not. But I definitely see a lot of people building very cool applications without a traditional ML background.

Matt Turck [10:53] And just to play it back, what you're seeing in the field is hybrid systems that combine foundation models and generative AI systems and traditional machine learning models. I'm repeating this because that seems like something that just about every practitioner in the field sees and agrees with. However, in the general public, there seems to be a narrative that generative AI is completely ripping and replacing all forms of AI that came before it. But just to confirm, that's not what you're seeing at all.

Chip Huyen [10:58] I would love to see people unseating XGBoost. I think it's, like, an uphill battle.

The Generative AI stack

Matt Turck [11:08] So the generative AI stack—that's the foundation of generative AI. What are the different components that people should know?

Chip Huyen [11:39] So when you look into building applications, I would think about maybe a development process, and maybe the stack should evolve to address your needs, right? So when you start with applications, maybe you start building, thinking about maybe you start with testing out the models. So you might want to do some prompt engineering and see how far you can get with good prompt engineering. Maybe you need to curate some evaluation metrics. You definitely need to evaluate, to design some evaluation metrics. So I think of this as the application development layer, right?

Chip Huyen [12:10] With prompt engineering, maybe with how to enforce structured output, security guardrails, definitely with evaluations. And then after you do it, so we've got application layer, and then we max out that performance there. And then we're going to, hey, maybe we need to change the model, right? Maybe we need to fine-tune the model. We need to make the model smaller, make it faster, inference optimizations. So that layer, when you actually make some changes to the model itself, is the model development or fine-tuning layer.

Chip Huyen [12:38] And then after that, I think you go into infrastructure. You deploy the applications, and now you have to scale up. You have to think about compute. You have to think about data storage. That's infrastructure. You start building out the platform so that you can make the deployment and iterations faster and more reliable. Basically, there are three—I think of them as three layers: application development layer on top, and then model development layer encompassing both developing a model from scratch and fine-tuning and making changes to the models.

Why are language models able to scale?

Chip Huyen [13:00] Ideally, you don't have to build a model from scratch. And then at the bottom, it's like infrastructure is powering everything.

Matt Turck [13:25] So scale, which you mentioned a second ago, seems to be—or is very much, I should say—at the heart of the entire generative AI approach. What is it that's so special about language models that makes them so reactive to this scaling approach that led to the ChatGPT moment?

Chip Huyen [13:50] Yeah, that is a really good question, and I feel like that's a question that I think today a lot of people take for granted. But it was not obvious before that language could be the way to scale intelligence, right? I think in the early days, at least when I got into AI—it was in 2014—people were still debating, like, would it be computer vision or language or reinforcement learning? Because computer vision was like, oh, because we developed ways to see way before we started developing language.

Chip Huyen [14:25] So maybe seeing is the way we scale intelligence, right? So it was not really obvious. And I remember back in 2017, I was at this OpenAI party, and somebody told me, like, "Hey, guess what? We just keep on throwing text, and now this model is pretty smart." So it feels like people just started to realize, like, okay, we just keep on getting more text, then we get much, much better models. And why language modeling? Because there were other text models, like machine translation.

Auto-regressive vs. masked models

Chip Huyen [14:54] I was actually very bullish on machine translation. I thought it was a really difficult task back in 2014, 2015. And now it was pretty much like people were saying machine translation is pretty much solved for major languages. I mean, we still have the long tail of lower-resource languages, but yeah. So the thing about language modeling is that it's a very simple task, and it's really elegant. So the idea is to get enough statistical information about the language so that you can predict what comes next in a sequence, right?

Chip Huyen [15:15] The idea is, I even say, like, hey, my favorite color is, then it should be able to predict that blue is going to be more likely than, say, car.

Matt Turck [15:20] And that's what you call autoregressive models, is that right?

Chip Huyen [15:51] Yeah, so I think that autoregressive is definitely one type of language model. So I think that's the idea of all language models, to encode statistical information. And that concept is actually not new. I think people employed that to decode, like, to break codes during World War II, right? People use that for games, and it's very, very interesting. So autoregressive, like you mentioned, is to predict what comes next, whereas the masked one is like you can have a context from both before and after and predict what is in the middle.

Chip Huyen [16:24] And for both kinds of language modeling tasks, the data is abundant. If you have some tasks like machine translation, you would have to curate, like, here is the original sentence and here's the translation, and it can be quite painful to curate that. But for language modeling, because you can just have any natural text online, there's so much of it online, you can just use it. And also, not just natural text, you can use programming languages, you can use codebases.

Supervised vs. unsupervised vs. self-supervised

Chip Huyen [16:46] And I think that it's just like, that's the nature of it. It's like you don't need to curate labels, like reference data, that you can use to train models that make language modeling so much easier to scale than other types of tasks.

Matt Turck [16:56] And maybe to put terms around it, can you quickly define for us supervised versus unsupervised versus self-supervised?

Chip Huyen [17:21] Yeah. So the approach of purposefully curating the labels for the models trained for fraud detection, right? For example, you have a transaction, here's a label, and the label can be fraud or not fraud. Or spam detection: the data is an email, and then the label is spam or not spam. So you have to curate, create those labels, the process of manually creating them.

Chip Huyen [17:54] And it teaches the model to learn from those labels. So that process is supervision. So the other spectrum is unsupervised. It's like you don't need to tell the model the labels, and the model figures it out. For example, for clustering. So if you throw in a lot of articles to the model and say, hey, try to group this into five groups. So you don't need to tell the model, okay, this group is technology or something like that. The model can do it.

Chip Huyen [18:23] A lot of clustering algorithms are unsupervised. Language modeling is somewhere in the middle. It's self-supervised. And the reason is that it still learns from some labels. For example, the next word in a sequence is a label that you need to learn to predict, but these labels come naturally. You don't need to manually curate it. You just get any text and you can generate a bunch of training samples. So yeah, so that's self-supervision.

Matt Turck [18:41] And still on the topic of scale, why does it matter how big a model is in terms of millions or billions of parameters? What difference does it make?

Why does model scale matter?

Chip Huyen [19:11] Oh, that is a very interesting question as well. I love how you're asking these very deep philosophical questions, and I feel like I really need a whiteboard to be able to explain all of this. So, people, if you don't understand me, trust me, my writing is better than my speaking. So why does it matter that the model should have a lot of parameters? So, the number of parameters usually approximates the model's learning capabilities. So it's just more like, with more—you can think of having more parameters as having more ways for the model to learn information.

Chip Huyen [19:50] So you can think—I'm trying really hard not to use a term like neurons—the brain having more synapses, you can learn more. It's not that equivalent. But yeah, so basically you can think of the number of parameters as more learning capabilities, or the learning capacity of the model. So, more parameters allow the model to learn more. So that's actually a very interesting question there. Is that like, why do larger models need more data to learn? Because the idea is that if the model has more capacity to learn, shouldn't it need less data to learn?

Chip Huyen [20:22] Does that make sense? If someone is smarter, right, it should learn faster from less data. Yeah, yeah, yeah. So I think this idea is like, because it has more capacity to learn, it could be a waste of this capacity if you don't teach it more. So yes, you can train a large model using a very, very small dataset, but that could be a waste of compute and a waste of that model's potential. You might achieve much better performance just training a smaller model with a smaller dataset.

Mixture of Experts

Chip Huyen [20:40] So, yes, larger models allow the model to learn more and give it more data, allow it to maximize its learning potential and be able to do much, much more powerful tasks.

Matt Turck [20:57] Could you go into some other approaches to make a smaller model very performant? In particular, I'm thinking of a mixture of experts. Can you maybe define for us what that is and what the general goal is?

Chip Huyen [21:27] So I do think the goal of making smaller models better is actually a very important goal. And it's just what everyone is trying to do. I want to point out that what is considered small or large is actually very time-dependent. What was considered large 10 years ago is considered tiny today. So I feel like what is considered large today might be considered small in the future. And we have seen time and time again, for example, the same LLaMA model families, right?

Chip Huyen [22:05] Like the LLaMA 3 models, the smaller model in the LLaMA 3 family probably performs better than the bigger model in the first LLaMA generation. So over time, we actually learn more about how to make models perform better at a smaller size. So, how to make smaller models better? I think there are a lot of ways. We can use better data. We've seen that higher-quality data can actually lead to better performance. People show that a lot.

Chip Huyen [22:37] Better training techniques, like new alignment techniques. Also, maybe different architectures that you mentioned, like mixture of experts. So mixture of experts is an interesting term because it has been reused for different meanings over the years, right? Before, we had expert systems, and mixture of experts meant different things from what people call mixture-of-experts models nowadays. So the idea is that, for example, you can have not quite human experts. You can have different—maybe you can divide the model into different experts, right?

Chip Huyen [23:22] And each expert specializes in some things. And then these experts, instead of training experts entirely from the beginning, which have a lot of parameters, you make these experts share some parameters, and then it has some kind of router in the middle to determine which expert is the most suitable. So the idea is these different components can share parameters to make it more efficient, parameter-efficient, where it can do multiple kinds of complicated tasks. Yeah, so I think that's definitely one pretty interesting approach.

Chip Huyen [23:58] But also, it's harder to train. So usually I don't see people saying, "Hey, I'm going to make a mixture-of-experts model today." People don't wake up and want to do that. It's pretty hard. I think for a lot of people, they could probably use something like, "Hey, maybe I do quantization," which universally works really well for a lot of tasks across models. Another thing people might try to do is just distillation. So you have a bigger model teaching a smaller model.

Chip Huyen [24:16] So that smaller model learns to mimic the behavior of the bigger model. So, yeah, I think that's a very, very, very fascinating question that you asked. And I might write another book about it.

Pre-training vs. post-training

Matt Turck [24:44] You heard it here first. Speaking of training, walk us through, in a nutshell, the different phases of how you train those models. There's pre-training, but in particular, I was very interested, reading the book, about the post-training phase, which is—there's a lot more to it than I had read about previously. So, maybe walk us through the steps, please.

Chip Huyen [25:15] Before I go into it, I hate the terms. Pre-training and post-training are a little bit confusing. I feel like the AI research community is really great at many things, but naming is probably not one of those. The process of training or creating a model like ChatGPT has a pre-training phase, which is when you train a model on the language modeling task. During this phase, the model gets really, really good at predicting what comes next.

Chip Huyen [25:51] It's like completions. You just say, "To be or not," it will complete with "to be." It's very good at that. However, people realized completions are good, but they're not very useful day to day because, let's say I ask, "How to make pizza," it might answer with "for 6" because it's trying to complete the sentence, like, "How to make pizza for 6," right? So a lot of times, completion is not always task solving. So that's where post-training comes in.

Chip Huyen [26:24] You teach this model, which has a lot of statistical information about all the knowledge of the world: now, how do you get it to respond in a way that is helpful to humans who are interacting with it? So in this phase, it's called post-training. People have multiple techniques, but what I'm going to say here is not always the only way to do it, but definitely a common way to do it. So in the first phase of post-training, maybe it's called supervised fine-tuning. You can do supervised fine-tuning.

Chip Huyen [26:59] So you can curate a bunch of, here's the instruction from humans, and here's how to complete that instruction. So if the instruction is, like, "Write me an essay about how wonderful my talk is," right? You can write an essay. So here's the response to the instruction. You train the model to mimic that human behavior of, here's instructions and here's how to respond. Then another phase that's pretty common is when you actually try to get the model to maximize the chance of it generating good responses and lower the chance of it generating bad responses.

Chip Huyen [27:38] So you can use techniques like reinforcement learning, like RLHF, reinforcement learning from human feedback, or DPO, direct preference optimization. Basically, there are a lot of other techniques around post-training. And unfortunately, a lot of labs that are doing it are not quite publishing papers about it. So a lot of work is just needing to know who to talk to, interview the right people, and try to get them to say it off the record.

Chip Huyen [27:51] That's very interesting.

Matt Turck [27:59] Is that because that's a big part of the secret sauce that those commercial proprietary labs want to preserve?

Chip Huyen [28:32] Yeah. So if I think about it, what makes a Claude model so different from ChatGPT and Gemini? In the pre-training phase, for a lot of companies, they have the same data because everyone is scraping the internet, right? Everyone is getting basically the same data. So for the language modeling task, everyone is optimizing for entropy, perplexity, right? So what makes these models really different? The post-training phase. So they curate data differently. They have different ways of collecting human preference and training for that.

Sampling

Chip Huyen [28:44] So I do think that post-training is what makes these really big lab models different.

Matt Turck [28:58] You mentioned a term called sampling in the book that is very interesting, and you mentioned it is very important in terms of understanding how those models behave. Can you maybe define that for us?

Chip Huyen [29:28] Sampling is really fascinating. I think writing sections on sampling is one of those things that brings me the most joy because I really like it. I feel like the topic is really underrated. So sampling is a process where a language model picks one output out of so many possible outputs, right? We talk about the language model encoding statistical information about language. So let's say I say the answer to this question is 70% yes, 20% no, and then 10% maybe, right?

Chip Huyen [30:02] So the model looks at all these possibilities: hmm, what should I pick next? So maybe 70% of the time it's going to pick yes, and 20% it's going to pick no, and then 10% maybe. So that is a sampling process. And the language model doesn't just sample one token or one word, right? It has to sample over and over and over again.

Chip Huyen [30:38] So sampling refers to different strategies to nudge the model to pick the output that is more valuable to you. So let's say, for example, one thing people do is use temperature. So you can nudge the model to pick more frequent tokens. For example, in the simplest way, you nudge the model to pick the most frequent, most likely token. So in that way, you would notice that the models become quite boring, because it would always pick what is most frequently spoken, the most common phrases.

Chip Huyen [31:14] First of all, you can ask, hey, what's your favorite color? People are like, my favorite color is blue, or my favorite color is red. It would speak very simply. It would be pretty unlikely to generate something like, my favorite color is the color of blue sky reflected in the still water, whatever, something creative. If you want something more creative, you would want to nudge the model to sample something that is less frequent. But now's the tricky part, because if you want the model to pick something really rare, it might become incoherent as well.

Chip Huyen [31:57] So basically, sampling refers to this whole family of different strategies to nudge the model to generate the responses most suitable for the task. And I do think this is fascinating and useful because it's a cheap way to improve an application's performance without having to retrain the model. It's also really useful for debugging applications. For example, maybe some model output is reasonable, but you can look at the probability. For example, if you ask a bunch of yes-or-no questions and the correct answer is yes, but the probability for yes is really low.

Evaluation as a key to AI adoption

Chip Huyen [32:14] So maybe the model is not that confident, right? So you want to look into that. So yeah, sampling is very cool.

Matt Turck [32:47] All right. So let's go into the general topic of evaluation, which you mentioned upfront is one of the key things in designing those AI systems. You actually have a sentence which I really liked in the book that summarizes it all, that says, “As teams rush to adopt AI, many quickly realize that the biggest hurdle to bringing AI applications to reality is evaluation. For some applications, figuring out evaluation can take up the majority of the development effort.” So you mentioned some of the specific challenges of why is it so hard to evaluate?

Matt Turck [32:56] How does one think about evaluating those models?

Chip Huyen [33:27] Yeah, evaluation is hard. I think I coined this term, what I call evaluation-driven development. This comes from engineering, the concept of test-driven development. The idea is that you develop applications that you can evaluate. Even though I think I see a lot of people excited about the latest marketing buzzwords, I think one thing I've realized from working with a lot of tech executives is that they're actually really smart. And I think, surprise, you become the SVP of this giant corporation because you're pretty smart.

Chip Huyen [34:08] But yeah, so I think a lot of business decisions are still made based on return on investment. So that's why it's really hard for people to double down on something if they can't say, hey, this is making real money for us. So it's not a surprise. It's not a coincidence that some of the most popular AI applications today are those that you can evaluate the output pretty clearly. So, for example, recommender systems. Everyone has a recommender system nowadays. And it's because with recommender systems, you can tell how much money it's bringing in by whether it's increasing, say, click-through rate or purchase-through rate, right?

Chip Huyen [34:56] So, you say, okay, after we launched the recommender system, now suddenly our purchase-through rate increased by 2 or 3%. And of course, you have to minus thinking about all the confounding factors, like campaigns, but maybe I'm talking very simplistically. Or for fraud detection, it's very common nowadays because you can tell very clearly how many fraudulent transactions your model should flag. And I stopped. For generative AI, one of the most common generative AI use cases today is coding. There are many reasons why coding is popular.

Chip Huyen [35:30] I'm going to say one of the reasons is that it's a lot easier to evaluate coding than others because here's generated code, right? You can evaluate: does it compile? And does this generate the expected output? Testing code is not new in software engineering. People have been doing all different kinds of tests, like unit tests, integration tests. So people know how to evaluate generated code. So yeah, coding is actually very important. So I do think that no matter how exciting a use case seems to be, if enterprises don't see a way to evaluate its outcome, it's very hard to nudge them to adopt it.

Chip Huyen [36:01] So I do think evaluation is the biggest bottleneck for AI adoption because unless we can develop a more reliable way to evaluate that application, that application is not going to get adopted. Or maybe it can. Maybe we need some billionaires to just steamroll fund it.

Entropy

Matt Turck [36:15] But so, what are the key concepts around evaluation? You mentioned a couple of those terms, like entropy and perplexity. What are the criteria? What are the methods? What should people know about?

Chip Huyen [36:48] Yeah, so you mentioned entropy and perplexity. Those are really fascinating concepts, and they actually guide the development of language models. Because most people today are not going to build a language model from scratch. You might do it for fun, but not at a scale where you can compete with OpenAI. I think entropy and perplexity are useful to know, but probably not what you are going to use day to day to evaluate your applications. I can talk about entropy and perplexity.

Chip Huyen [37:20] I think they're really, really, really cool concepts. One thing I want to mention about that is, we want entropy to be lower, to make things basically more predictable. So, for example, if the model is getting really good at predicting the next token, that means the training data becomes— that language becomes more predictable to the model, right? So the entropy is going to be low. And over time, people found out, hey, if I could just decrease entropy somehow, users are happier.

Chip Huyen [37:56] All the users using the applications become happy. And then the question is, how far can you go? How low can the entropy go? Because absolutely, it can't go to zero. I think people have been talking about whether there is a lower bound of how far you can go with entropy. How much room do we have left to push the performance of these language models? And there's this concept of irreducible loss. So language has some certain aspect of unpredictability, right?

Chip Huyen [38:31] There's just no— I don't think we'd ever reach the point where we can predict the next token perfectly, because there's always some variation in the way we speak, right? So I do think there is this irreducible loss. And I'm not sure if you saw a bunch of people talking recently about the end of pre-training. And I think there are multiple reasons. One is, we don't have data for pre-training. The second is just that the perplexity, the entropy of this language, is pretty, pretty low.

Chip Huyen [39:06] And it might be very, very close to what could be theoretically possible. When Claude Shannon introduced the concept of entropy, he did some pretty fun exercises. He asked a question like, hmm, what is the entropy of the English language? So that could be the lower bound. But because he did that back in the 1950s, he did that based on a very, very short sequence of maybe 10 words. And the interesting thing about entropy is that the more preceding sequence, the longer the context, the easier it is to predict the next token, right?

Chip Huyen [39:43] Usually, if you just tell me one word, it'll be very hard for me to predict the next one. If you just say, "I," right, I could go with "I am," "I want," "I love," "I hate," right? But if you give me a pretty long sequence, for example, even if you say, "Today, I would like to welcome my—" I can predict "guest." The next word is going to be "guest." So the longer the sequence, the easier and more predictable the next token is, and the lower the entropy.

Chip Huyen [40:04] So Claude Shannon did that exercise in the 1950s with a very short preceding sequence. So I would really love to see if somebody today does that study, but for really, really, really long sequences.

Evaluating AI systems

Matt Turck [40:35] And so, if a concept like entropy is not something that people who build those AI systems in real life—not the model developers, but the AI engineers who deploy AI systems—need to worry about or use in their evaluations, what should they use? What are some of the key techniques and key concepts to evaluate AI systems in production, in real life?

Chip Huyen [41:07] Yeah, so I think, like a lot of corporations, it should make money. I think the ultimate metric is whether it's making you money or not. But I feel like, exactly like you brought up, many things can cause whether a company makes money or not. So I think, for applications, it's really, really important to understand the use cases well so they can design the set of metrics. Then you can work backward from that and map it to the model metrics that you care about.

Chip Huyen [41:27] Let's say, for example, you do a text-to-SQL model. I feel like back in 2023, every week you would get some engineers like, "Hey, check out my new text-to-SQL model."

Matt Turck [41:27] Pretty much.

Chip Huyen [41:54] I was like, wow, wow, people would do anything to avoid writing SQL queries. Let's say you're a data company, right? You usually interact with your data using SQL. It was like, "Oh my God, writing SQL is so painful. Let's help people use natural language to write that." Let's say you have a text-to-SQL model. Then how do you know that the model is good? So maybe you can start thinking from the user perspective, from your perspective: why do you want to develop this model?

Chip Huyen [42:27] First, maybe you want to improve users' productivity. So maybe you can use the metric of time to completion. So maybe before, without our tool, users would take on average maybe three minutes to write a SQL query, but now with the tool, it takes only one minute. Having this metric would be very useful. Or for customer support, you can also similarly, for example, before you can respond to users, it should take two hours for the agent to get to the users, but now we can respond instantly.

Chip Huyen [43:06] But that's not always the case because by default, if you respond to it automatically, it will always be fast. So I need to think more about the case of, are users happy? So that is the ultimate case evaluation metric. I do think it's really dependent on the use case and your company and what you care about. But then you work backward from that. It's like, okay, now I don't want to deploy the application yet because I want some validation offline to be able to know whether this is good or not.

AI as a judge

Chip Huyen [43:22] So you need to evaluate. You can create evaluation systems to evaluate that.

Matt Turck [43:37] Still on the topic of evaluation, an interesting tidbit you talk about is the concept of AI as a judge. So somebody or something needs to evaluate the system. It could be a human, it could be AI. What are your thoughts on the pros and cons of AI as a judge?

Chip Huyen [44:08] Yeah, I do think AI as a judge is a very promising approach. I think when ChatGPT first came out, AI as a judge was just like, AI is not reliable enough to be entrusted with that crucial task. But I think nowadays, you talk to teams, I think most teams have some variation of AI as a judge going on. So AI as a judge is pretty interesting. The idea is that you have an AI evaluating the outputs of other AI, and it's especially useful in production.

Chip Huyen [44:48] The idea is that, let's say you use a model to generate a response. A lot of people were like, "Oh my God, what if the response is not safe? What if the response is crazy? What if it's saying something like, 'Gimme sue'?" So maybe you can have another model just to double-check this one and give a score and send it back. So AI as a judge has been shown to work pretty strongly correlated with human judgment. And the tricky thing about AI as a judge is that AI as a judge is not as objective as other metrics like F1 score.

Chip Huyen [45:30] So what that means is that when somebody says F1 score, you know what that means, you know how it's defined, right? And if I run a calculation for F1 score again, even using my own F1 score code, I will get the same F1 score, ideally. But for AI as a judge, it really depends on who the judge is, which model is the judge model, and what the prompt is. One thing I noticed is that a lot of those judges can evolve over time. With evaluations, ideally you want the evaluation method to be stationary so that it can benchmark your application over time.

Chip Huyen [46:12] Let's say that yesterday, the evaluation metric was maybe 90%, and today it's like 92%. And you know that, okay, so my application is getting better. But with AI as a judge, what could happen is that the judge itself changes. So it's not comparable between 90% and 92%. And I once talked with a team at a pretty big company. It's a pretty common scenario with a lot of companies, especially when they're bigger, is that they may have a team that developed AI judges, like maybe write a prompt for an AI judge.

Chip Huyen [46:46] Maybe the judge could be like a faithfulness score or a relevance score. And downstream teams should use the judge. So this one engineer came to me and said, "Hey, we have this faithfulness score of 90%." And I was like, "Okay, that's great. So what is the prompt you use for the judge?" And he was like, "I actually don't know. I just use this off the shelf." So I do think it's really, really tricky when you don't have control over the judge.

Why prompt engineering is underrated

Matt Turck [47:14] So we were just talking about prompts. So let's turn to prompt engineering. And you wrote that prompt engineering's ease of use can mislead people into thinking that there's not much to it. So maybe a quick reminder on what prompt engineering is in the first place, and how should AI engineers think about it or approach it?

Chip Huyen [47:46] Yeah, yeah. I think when I mentioned that I have a section on prompt engineering, I did have a few people roll their eyes, like, "Oh my God, prompt engineering, blah, blah." A lot of people don't take prompt engineering seriously because they think there's not much engineering to it. So maybe it goes back to, what is a prompt? A prompt is how you communicate with a model.

Chip Huyen [48:18] So I think of writing prompts as just like writing. Does that make sense? You can think of writing prompts as human-to-computer communication. And just as with human-to-human communication, anyone can do it, but not many people can do so effectively. So yeah, because anyone can say, okay, I can just write this prompt, people think, okay, if anyone can do that, it's so easy. There's nothing to it. And especially early on, it was quite misleading when we had a lot of hackiness when it came to writing prompts.

Chip Huyen [48:54] For example, I think one of the funniest tips I saw on prompt engineering is, if you tell the model, "Answer correctly, and I will give you $200." Bribing the model. Yes, bribing the model. So it was like, okay, you just do stuff like that and get more things done. But even though there's a lot of tweaking the instructions to get what you want, I do think that it can be very systematic.

Chip Huyen [49:26] You need to make it very systematic. So if you consider each prompt an experiment, there should be versions of prompts. You should be able to systematically track your progress with different prompts. You don't want to just use a prompt and have somebody make random changes, have no idea what's going on, and downstream people have no idea what changed about applications and why the output is different. I do think prompting is very easy to get started, but to do so effectively does require a lot of practice and a lot of discipline to do so systematically.

In-context learning

Matt Turck [49:42] What is in-context learning when it comes to prompt engineering?

Chip Huyen [49:47] By the way, do you think most of the audience would know what in-context learning is?

Matt Turck [49:49] No, but they're gonna learn thanks to you.

Chip Huyen [50:18] Okay, so in-context learning is actually pretty—nowadays, it's one of those things people take for granted. But it was a pretty novel idea when it came out. So, in-context learning: now, when you talk to a model, you give instructions, you give some information. A lot of people call the entire thing that you input into the model to get the model to do what you want the context. So, the context is actually very confusing.

Chip Huyen [50:45] We're going to go into that later. But yeah, so you give the model a bunch of information to get the model to do what you want. So that is a context. And if you give it some examples, like if I say car, you say vehicle. If I say bananas, you say fruit. If I say house, you say building, right? And it says this is the next thing. Maybe if you say train, it will know how to output vehicle, for example.

Chip Huyen [51:19] I think when you put the examples into the model, the model is able to learn from these examples and output the correct category or the correct output. And that idea was novel when it came out because before, if you wanted the model to be able to predict, to guess the correct output, we had to train the model especially for that. Like before, you want to say, oh, predict the category of this object, right? We need to curate the training data of name to category, had to train the model on it.

Few-shot learning and zero-shot learning

Chip Huyen [51:54] But now with language models, it's just like, it depends. You don't need to train the model from scratch. You just need to input some examples in that, and then it gets it to exhibit the behavior that you want. So that is in-context learning, like the model learns from the context given to it. I have this whole term of few-shot learning, zero-shot learning. Zero-shot learning is when it can do what you want without any examples at all.

Chip Huyen [52:23] For example, it can say, "Tell me whether this email is spam or not spam," and give it the email and no other examples. That's zero-shot learning. But then you have maybe five emails, each of them with a label of spam or not spam, and then the email you want to classify. So now you have five examples for the model to learn from, and now it's five-shot learning.

Matt Turck [52:30] So it's a really big deal, right? Because, as you said, you don't have to go back and retrain a whole model with brand-new data.

Defensive prompt engineering

Chip Huyen [52:57] You can extend an existing model with knowledge without any kind of coding or training. Yeah, it's a big deal because it makes language models so versatile and so useful in a lot of applications. So it's like what makes a general model. Before, if you want a model for a task, you need to train a model for that task. But now you have a model trained generally for language modeling, and now you can just adapt it to any task with in-context learning.

Matt Turck [53:21] And there's a part in the prompt engineering section that I thought was particularly interesting around defensive prompt engineering, so defense against jailbreaking, information extraction, that kind of thing. Can you talk about that and maybe what we have learned about making those AI systems resistant to those kinds of attacks?

Chip Huyen [53:50] I do think that's a topic that's getting increasingly important, especially as AI is being used for more high-stakes tasks and more complex tasks. And the second is that now AI has increasing access to more tools and it can make changes. So, we do want AI users to be safe, protecting not just users but also developers of those models, that nobody gets sued. So I think that it's also one reason that makes a lot of people go to proprietary models instead of open-source models.

Chip Huyen [54:34] Because I'd say if you use a model developed by a company through their API, this company is responsible for putting guardrails to make sure that model behaves safely. It doesn't say anything racist, sexist. Somebody asks this question about praising Hitler, maybe you can say, no, I'm not going to do it. There's a lot of guardrails around this. But if you use an open-source model and you deploy it yourself, you're responsible for it yourself. So I do think it's like, of course, open-source model developers, they try their best to make the models safe as well.

Chip Huyen [55:05] But they also have less visibility into how the open-source models are being used, which gives them less information for them to make the model safe. So I do think defensive prompt engineering is very important. It refers to the process of writing prompts in a way that makes the model safe. So, for example, maybe nobody actually knows what the ChatGPT system prompt or Claude system prompt is, but I'd say it contains a lot of language that just tells the model, do not respond to this kind of request.

Chip Huyen [55:28] If you are asked this, say this. Do not be this. Do not be that. Do not be that. Right.

User prompt vs. system prompt

Matt Turck [55:53] And just maybe to make sure everyone understands. So there's a user prompt, but there's a concept of a system prompt, which is like the prompt behind the scenes, right? Like the prompt behind the prompt that tells, once you as a ChatGPT user have typed in something, then there is a second layer that tells the model to do and not do certain things. Is that fair?

Chip Huyen [56:19] Yes. Matt, I think you could be a great teacher. That's because I read your book. Thank you. Yeah. So you do have the user prompt and the system prompt. So the system prompt is like, say I'm an application developer and I develop an app of, like, hey, given disclosures for a house, users can ask questions about the disclosures. So as an application developer, I write a system prompt like, hey, model, act like a real estate agent, and given disclosures that users are going to ask you, act professionally, be nice, be kind, be brief.

Why RAG is here to stay

Chip Huyen [57:07] And then user prompts are like, "Here's a disclosure. Tell me, how big is a lot? How is that compared to any noise complaints?" System prompt is created by the application developer. User prompt is when users interact with the applications via whatever language they use. System prompt is when you need to put in all these languages to make sure that the application acts in a way you want it to act.

Matt Turck [57:42] All right, let's spend a couple of minutes on another very important part of AI application architecture, which is RAG. Maybe a quick reminder for folks about what RAG is. And then one specific question: the topic sentence that caught my attention is when you wrote that many people think that a sufficiently long context or context window will be the end of RAG, but you don't think so. So what do you mean by that?

Chip Huyen [58:08] Yeah, so I think originally RAG was developed to get around shorter context, right? So I think originally it was introduced in 2020 or something, maybe, in a research paper. But yes, the idea is, like, for a lot of knowledge-based, knowledge-intensive applications or questions, when you cannot fit the entire knowledge needed in the context, then you need to retrieve it from somewhere. For example, if you need to rely on Wikipedia and you cannot fit the entire Wikipedia in your context, so maybe you need to find the article most relevant to the question, retrieve it, and put it in the context instead.

Chip Huyen [58:43] So that's the original idea. It was designed to get around that issue. And then some people were like, okay, as the models get really long context, maybe you don't need RAG anymore because you can just dump the entire enterprise database into the context and you can just retrieve it. So I do think the question of whether long context will make RAG obsolete is similar to: will larger computer, like laptop memory, make data centers obsolete?

Chip Huyen [59:14] So I feel like no matter how big my phone memory or my laptop memory is, I will always run out of memory. So people, we're only having more and more information. So it's like, I think we always expand our usage to fit in whatever context length that's going to be available. And the second thing is that— I think the second reason may be actually more important, at least for now, is that just because a model can fit in a million-token context doesn't mean that it can process that million tokens efficiently.

Chip Huyen [1:00:02] I think this, for example, I was actually using Claude and ChatGPT for some of the novel writing. I actually really want to write a math novel. What I found out is that because I want to write stories, you need to be able to keep track of events in the past. And I found anything that's over 10,000 tokens long, the model would just completely forget, confuse timelines. It would think that this character actually met that character instead. It is really not efficient at all at understanding large context.

Defining AI agents

Chip Huyen [1:00:31] And I feel like some model developers, when they have the RAG guidelines, they have some guidelines: hey, for anything beyond these X things, maybe try to use RAG. If it's shorter than that, you can dump everything into it. But really, it really depends on the use case. And I think people really need to test out the context efficiency for their application.

Matt Turck [1:00:49] All right. So, as promised, let's close with everyone's favorite topic of the day: AI agents. So I guess let's define the term, first of all, because everyone seems to be using different definitions for the term. So what is an AI agent for you?

Chip Huyen [1:01:01] That is—I feel like it's a trap question. You were like, okay, everyone disagrees on the term. Now tell me what the term is. By the way, Matt, do you invest in any agent companies?

Matt Turck [1:01:10] Of course. I'm a VC. I have to. The two things that I need to do as a VC is, one, I need to have a podcast. Two, I need to invest in AI agents. Those are the rules.

Chip Huyen [1:01:39] So I think it's like, I actually like a lot of my early reviewers. They have very, very, very extreme bipolar opinions on agents. So I decided, okay, for this case, agent is not a new term. It's not like a term that no one has ever heard of. It's actually a term used in AI for a very, very long time. So I was like, okay, let's just go back to the basics. And I just took out some textbooks from the '80s and '90s and see how people define agents.

Chip Huyen [1:02:19] It actually makes things a lot easier. So I actually base my definition in my book based on Peter Norvig and Russell's book from the '90s. It's a really good book. So basically, they define agents as anything that can perceive the environment and interact with the environment. It has two components, like the sensors to get information from the environment, and then it can perform actions in that environment. What does it mean in the context of AI-powered agents? AI-powered agents already can interact with the environment.

Chip Huyen [1:02:55] For example, if you are saying, hey, search the internet, that means it's retrieving the information from the internet. And if it says, hey, send an email, then that means acting on the environment by sending out information. So AI agents are basically defined first by the environment they operate in. So let's say ChatGPT operates on the internet, then the environment is the internet. If you have an agent in Gmail, then Gmail is that environment. If you have a coding assistant agent, then maybe whatever coding editor that you use, like VS Code or whatever, or terminal is the environment.

Chip Huyen [1:03:14] It's characterized as the environment it operates in. And then, given a user's task, this agent leverages tools.

Matt Turck [1:03:17] Leverages tools.

Chip Huyen [1:03:45] It has access to perform that task. And how does it do that? It would need some type of brain that determines, hey, what task is this? What tool should I use? How do I invoke the executions? And how do you determine that this task has been accomplished? So in terms of the AI-powered agent, that brain is an LLM, like it's a model. So we have GPT-4-powered agents or Claude-powered agents. So if you think about AI-powered agents, you basically think about, hey, what are the tools that it has access to?

AI agent planning

Chip Huyen [1:04:04] And second, how good is its planning capability? Because other aspects, like the environment, are pretty much defined by the agent developer and the user, and the task is supplied by the user.

Matt Turck [1:04:15] And then there's a central concept to agents, which is planning. How do you think about it, and what is specific about planning from a model perspective?

Chip Huyen [1:04:45] Planning is very, very, very hard. So I think we basically talk about an agent having two big components that we need to work on: tool use, right, and the different set of tools the agent has access to; and planning, how to use the set of tools efficiently to solve the given task. Tool use, actually, in the early days, we already had function calling. So, what is tool use?

Chip Huyen [1:05:18] Invoking a tool by calling a function, supplying the parameters to do the function. So function calling, we pretty much understand a lot more about it. But planning, how to use those tools effectively, how to come up with a roadmap, an outline of a plan to solve the task, is really hard. So even something very simple can require a lot of reasoning steps and a sequence of actions. Planning is also not a new problem. Maybe what is new is using LLMs to solve planning, but planning itself is not new.

Chip Huyen [1:05:52] Planning, at its core, you can think of it as a search problem. What that means is, given a task, like, hey, you have a goal over there, there are many, many different paths toward that goal. Maybe you need to turn left first, turn right first. There are many different ways. Planning is basically, you search through all the possible paths and choose the one that is the best path. Or if there's no path, you need to determine that there's no path, there's no possible way to solve this task.

Chip Huyen [1:06:11] Planning is a search problem. I think in many AI textbooks, even the books I mentioned from Norvig and Russell, they have a huge section on planning. The challenge for LLMs for planning is that I feel like it can probably go into multiple minutes here, but I do think planning is—you need to be able to understand, not just generate a sequence, not just know, not just say what to do next, but you should be able to predict what could be the outcome of that.

Chip Huyen [1:07:01] Does that make sense? For me, it's really hard to decide, should I do A or B, if I don't know what is the output of that? What's the expected outcome? So I do think there are certain techniques that make LLMs better at planning. So, for example, before trying to do an action, try to predict: if I do that action, what would happen? If you turn left and turning left is a cliff, you probably don't want to do it, right?

Chip Huyen [1:07:33] Planning is incredibly hard. And I do think in my book, I have a very long section on planning. I actually made it public on my website, the section on planning. If what I'm saying is a little bit too abstract, go read the blog post. It's free and it's online. Basically, the idea is: can LLMs plan? Is there any fundamental reason why LLMs cannot plan? Because we have people on one side of the spectrum, like Yann LeCun from Meta, who's an incredible scientist, but he also said that autoregressive LLMs just cannot plan, and they kind of disagree.

Chip Huyen [1:08:17] I think maybe we just don't give LLMs enough tools to be able to plan effectively, or maybe we just don't have models good enough—maybe stronger models become better at planning. So I think that section discusses that and discusses different tips, like how to get the models to plan more effectively. For example, to plan well, you need to understand the set of tools you have access to pretty well. If the tools are too confusing to use, that's actually very hard for the model to efficiently plan.

Training data as a bottleneck to agent planning

Chip Huyen [1:08:51] You need to look at the toolset the model has. Maybe you need to change the name to make it more understandable, write better documentation, or maybe break the tool into simpler tasks. One thing I want to say about planning: it's actually not obvious, and I find it quite fascinating. It's actually hard for us to create data to train models to become better at planning. And the reason is that when we ask humans to generate what they consider the best plan of action for a task, it's actually not quite the best plan for AI.

Chip Huyen [1:09:33] Because what is easy or efficient for humans is not the same as easy and efficient for AI. For example, browsing 1,000 websites and summarizing 1,000 websites would be really boring and slow and tedious for humans. As a human, I probably would not do it. Actually, I did it for my—I tracked 1,000 repos. I did it at one point, but it's a very tedious and painful task. But for AI, it's actually really easy. It can just browse 1,000 websites and generate summaries at the same time.

Chip Huyen [1:09:58] One challenge in generating data to train the models to do better planning is that we can't quite rely on human labelers or annotators to do so. There's a whole school of thought: how do we come up with, how do we generate good plans for AI to learn from so it can learn to plan better?

Matt Turck [1:10:32] All right. So it's been wonderful. There's more chapters, more topics. We could talk about dataset engineering. We could talk about inference optimization and all the things. But hopefully that gave listeners a good flavor for what you discuss in this book, which, again, is amazing, and I very much truly enjoyed and fully recommend to anyone interested in the topic of AI in general and AI engineering in particular. So the book, which is both in electronic format and starting to ship in physical copy—my physical copy is going to arrive soon, I'm told.

Matt Turck [1:10:50] There is a GitHub repo associated with the book as well. Is that right?

Chip Huyen [1:11:14] Yeah. So I think in the process of writing the book, I went through so many resources. I think the book itself references over 1,000 links. And I myself personally went through a lot more links. So I found a lot of it was very helpful. So the GitHub repo has about 100 of the resources I found most helpful in the process of writing the book, so that if you just go through the resources directly, I think there are a lot of great learnings there from everyone.

Chip Huyen [1:11:50] And so, the repo also has a table of contents and summaries for each chapter and some prompt examples and stuff like that. Yeah. This is not a tutorial book, so there are no coding examples. It's pretty—I hope it's a good thing, because I felt like all the frameworks today change so fast. So I feel like any coding example using any of them is going to go outdated pretty quickly.

Matt Turck [1:12:06] Wonderful. Thank you so much. I think this book is going to be a major hit. Really appreciate you coming on the pod and telling us all about it and sharing some of the key insights. Really appreciate it. Really enjoyed it. Thank you so much.

Chip Huyen [1:12:07] Thank you.

Matt Turck [1:12:28] Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.