Humanloop: LLM Collaboration and Optimization with CEO Raza Habib
The MAD Podcast with Matt Turck · with Raza Habib, CEO, Humanloop
Raza Habib is the CEO at Humanloop. We cover why subjective, stochastic LLM outputs require evaluation across development, production, and regression testing, how tying feedback and observability to prompt editing cuts fixes from lengthy deployment workflows to the same day, and why teams often prototype with GPT-4 before fine-tuning smaller models for cost and latency.
Chapters
Transcript
Full episode
Matt Turck [0:47] Welcome to The MAD Podcast. Let's jump right into it. What's the elevator pitch for Humanloop?
Raza Habib [1:07] Yeah, great question. So, at its core, Humanloop's helping product and engineering teams build reliable applications on top of large language models like GPT-4. So we help them find, manage, and version prompts for these applications, optimize them, and then measure how well they're working. So, for example, Duolingo uses us to develop a lot of the prompts behind their AI-focused features. And then Humanloop becomes both a development environment and a CMS for their AI prompts over time.
Raza Habib [1:37] A lot of companies building LLM apps today don't have established workflows for this. Everything's really new. They've got domain experts doing prompt engineering. They've got software engineers who are actually implementing things. It's shared in Slack. It's very hard to evaluate. There isn't an established set of tooling like there is for more traditional machine learning and software. And so we're building the pieces that are needed to version, manage, and evaluate these components.
Matt Turck [1:57] Okay, cool. So, for people like me, I do this annual MAD Landscape where I have a lot of logos, a lot of categories. So where do you put Humanloop? Are you MLOps? Are you LLMOps? Are you prompt engineering? Are you all of the above? Or maybe it's a futile exercise and there's just not enough categories.
Raza Habib [2:15] Some of these categories are somewhat overlapping, and also there's a lot emerging. I think that I would describe us probably as maybe the first LLMOps company that there was. And over time, we've actually narrowed our scope as more players have come into the market. We haven't had to do everything. And so there's two pillars that we really focus on. And one is this collaborative development environment where non-technical people, product managers, and domain experts collaborate with the engineers on prompt engineering, coupled very closely to tools for evaluation and monitoring.
Raza Habib [2:36] So those are the two things we're focused on. It's LLMOps, I guess, at its core, but those are the pillars. And I think they belong very closely together for many reasons I can go into as we discuss.
Matt Turck [2:58] Okay, great. So what would you say are the specific challenges when evaluating and monitoring LLMs? Obviously, the category of evaluation and monitoring in the software world has yielded huge companies like Datadog in particular, but that's one world. How is the world of LLMs different? What are the specific challenges and opportunities?
Raza Habib [3:21] Yeah. So machine learning differs from traditional software in a couple of key ways. And then generative AI, as a kind of subset of machine learning, differs even further. So the first transition when you go from traditional code to machine learning is that it's no longer deterministic, right? We're used to, for software engineers, writing a program, you run it, you get the same results each time. You can write a deterministic test. And people are doing performance monitoring with something like Datadog, but they're not expecting that when they run the code each time, they're going to get different outputs.
Raza Habib [3:53] Once you move to the world of machine learning, now it's stochastic. And not only that, but you're now specifying what the program does via a dataset and a training process rather than deterministically in code. And so evaluation in machine learning focuses on accuracy metrics and things like that. And when we go to LLMs and generative AI, we go one step further, where the use cases that people are applying these to are very general and often very subjective. To pick a couple of examples, if you're helping someone draft a sales email or you're writing marketing copy, there isn't any longer a kind of ground truth answer that you can compare against for the model to know whether it's correct or not.
Raza Habib [4:30] Even if you're doing question answering where someone asks a question, there are many, maybe many, many different ways to correctly answer that question, with different wordings and different ways of expressing what you're saying. And so some of the challenge comes from the fact that it's so much more subjective and it's stochastic, so that it becomes difficult to know what does good look like, and then to measure that performance once you're in production. And so evaluation is important, I think, for LLMs at three stages.
Raza Habib [4:55] And each of them is challenging in different ways from traditional software. So during development, you have to make a lot of different design decisions. You're choosing between different prompts, different models. People are building retrieval-augmented systems. So they have to choose over the different components of an information retrieval system. And so there's this combinatorially large set of different things we're choosing between. And if you can't measure what good is, then it's really easy to just keep changing things and not knowing, like, am I getting better or am I moving closer to my goal?
Raza Habib [5:25] So during development, being able to have some quantitative feedback on, like, what does good look like, is very important. Then you deploy to production, and now the range of inputs that your users put through the system is really broader than maybe what you could have done during development and testing. And there's often surprises. How do you know that the system is continuing to behave the way you expected? It's not making things up. It's not saying things that are embarrassing, and it's just giving users a good experience.
Raza Habib [5:54] So there's a monitoring piece that is different to what you can do with traditional monitoring tools. And then finally, there's regression testing, which is, okay, OpenAI's model changed, or I did a new fine-tune, I changed a prompt, the vector database changed, I'm deploying something new. How do I know that it hasn't gotten worse on the core use cases that I care about? And if you don't have good evaluation and monitoring in place, you can't do that either. So those are the three areas where it comes up.
Raza Habib [6:03] And for each of those, it's really different from traditional software because of the subjectivity and the stochasticity. Great. Okay.
Matt Turck [6:10] So how are you solving the problem? Maybe taking those three in turn, what does Humanloop do to address the issue?
Raza Habib [6:33] Yeah. So we provide the ability to get evaluation data, both from human feedback in various ways and also in an automated way. So I'll maybe take the human feedback version first. And this is actually what the first version of Humanloop's product started out doing, which is that because it's very subjective, there's some sense in which the only real ground truth is your customer satisfaction or opinion of the experience they get out of the product. And so we make it really easy for people to instrument their applications with the ability to capture feedback.
Raza Habib [7:02] So the simplest version of this is things like thumbs up, thumbs down that you might have seen in many applications. You see it in ChatGPT, but people also tend to collect implicit signals of user satisfaction. So, the actions that they take after interacting with a particular LLM app. And also, if people are able to edit any generated text or generated content, then capturing those edits can be a very useful signal of how well things are working. And then in the Humanloop app, we triangulate those sources of feedback back against, okay, what model created it?
Raza Habib [7:26] What inputs created it? And we give product teams the ability to dig into that data, analyze it, understand what's working well and what isn't, and why. And then, critically, to also, in the same application, take actions to change things to make them better, to edit prompts or to fine-tune models. So the human feedback component is one big part of it. And that feedback is typically either coming from end users, or sometimes it'll come during development from an internal annotation team that maybe later will be switched to production feedback.
Raza Habib [8:02] And then we also have the ability to do automated evaluations, where these are evaluations for more conventional metrics than just traditional code. There are some metrics that people still calculate. They want to measure things like latency, maybe for things where there are scores, like ROUGE and BLEU, that measure token overlap with a kind of ground-truth test set. So things you would have in a traditional MLOps software, those are available. But then in addition, what you wouldn't normally have is model-based evaluations, where you actually have another LLM that is defined as an evaluator and reads the inputs and outputs of what's being given to the model.
Raza Habib [8:33] And then scores them in a particular way. A lot of people have been trying to do this. I'd love to go into a little bit more depth if we have time, because I think that there are bad ways and good ways to do this. And we're trying to basically productize it in the app such that you're naturally on guardrails and you set this up in a way that's going to work well.
Matt Turck [8:36] Yeah, let's do it. Let's double-click on this.
Raza Habib [8:37] Yeah.
Matt Turck [8:38] Super interesting topic.
Raza Habib [8:53] So I think when you come to use an LLM as part of the evaluation piece and you're trying to get it to score how well things are working, the danger is you're now adding in a second form of uncertainty and noise. The evaluator itself can have some noise in it. And so—
Matt Turck [8:56] Then you need an evaluator to evaluate the evaluator.
Raza Habib [9:09] Exactly. It's turtles all the way down. So in order to avoid that, we try and put some guardrails and constraints over how you create these. LLMs are pretty good at answering simple factual questions about inputs and outputs. So, for example, if you have a system that's doing question answering where it first retrieves from a database and then uses that retrieved context to answer the question, if you ask the model, hey, was the final answer based on the retrieved context?
Raza Habib [9:40] Yes or no? That's an answer that a model will get right very consistently, right? It's a factual question based on very simple reasoning, looking at input and output, and it's Boolean in nature: yes or no. But that gives you a really strong signal: do I have hallucination problems with my application, for example? And so we restrict people to generating evaluators that either have these binary yes-or-no-style questions, looking at the inputs and outputs, or we get them to return a numerical score.
Raza Habib [10:10] One of the mistakes that I have seen people make in the past is they'll do things like ask an LLM to rank or compare various outputs. And although this is getting better, there's some evidence to suggest that the order in which you ask the question affects the LLM's output. And so it's very easy to get very biased results. And so we make sure that that doesn't happen. Being very careful about what we allow users to define as LLM evaluators ensures that they work reliably.
Matt Turck [10:39] Yeah. And hopefully not an unfair question, because Datadog has built a multi-billion-dollar business, but let's assume Humanloop finds regression or different results from the same prompt. Then what can you do as an enterprise? Obviously, knowing that is an issue is essential, but is there a way you can fix that, or should you just be aware?
Raza Habib [10:59] No, I think this is one of the strongest arguments for why a new set of tools is needed and why using existing platforms for monitoring and observability is less productive. And it's because with generative AI, the speed with which you can make interventions is extremely high. And having that in one combined platform is actually one of the powers of this. So you're absolutely right. Within Humanloop, I mentioned we have this kind of interactive environment, both for prompt engineering, and we also have the ability for people to fine-tune models, which is where you do a little bit of extra training on a new dataset.
Raza Habib [11:36] And so what will often happen is people will find a bug. They'll basically be exploring the log data within Humanloop. They'll find an issue. They'll reopen those data points back into that interactive environment, where they can now run what-if-style analysis. So they can change the prompt or they can change the information retrieval system a little bit and see what the impact was. If they're able to then fix that issue, they then run a regression test. They say, okay, is this new prompt still performing well on what worked before?
Raza Habib [11:59] And if the answer is yes, they can actually promote it straight to development or production from within that system. And so actually, a product person or a domain expert who's able to go in and find a bug is also often able to do the intervention right there. And the turnaround time can be sometimes minutes to hours, but it's usually within the same day. And I think that if you don't have a system like this, then your logs typically live in a completely different system.
Raza Habib [12:28] So you find the issue, but now you need to say, okay, which model or which prompt was responsible for this issue? Okay, where is that stored? Okay, it's in code somewhere. Now get the prompt, put it back into some kind of interactive environment, figure out how to change it, and now go through a whole deployment process to get it back in production. That can be really slow and actually is unnecessarily slow. This is an opportunity of LLMs. Before, with machine learning, if you wanted to update a model, it was actually quite cumbersome.
Raza Habib [12:47] It required a full retraining run. And so you couldn't have as quick an intervention. But because prompt engineering allows you to intervene really quickly, I think it's very important to have evaluation and observability very closely connected to the prompt engineering tools.
Matt Turck [13:14] And the collaborative aspect around prompts, is it prompt engineering as in engineering inputs to help shape the model direction in a certain way, or is it like end-user, business-user prompts where you have a library of different prompts? So you say, well, if you want to get the best result for this marketing query, this is how we do it at Pfizer or Bank of America or any customer?
Raza Habib [13:37] So it's much more the first one. So people are building an end-user-facing application and they're trying to customize a very general-purpose model to that specific use case. And so in some ways it's just like writing code, but the code now happens to be in natural language and is really specifying to the model what you want it to be able to do. If we try and pick a concrete example, maybe you're trying to do a summary of a call for a salesperson. Well, the kind of summary you want for a salesperson would be very different for a different audience.
Raza Habib [14:04] And so it's explaining to the model, hey, this is for a salesperson. I would like you to be able to extract the budget and who the decision-makers are and what their biggest needs were and when the next stage of actions are. And really, it's getting to the point where prompt engineering is much more like writing the spec of software. It's really just about articulating very clearly what you want the model to do. And that's something that's often best done by product people or domain experts.
Raza Habib [14:29] And so one of the things that I'm most excited about is the extent to which this actually democratizes access to people who are closer to the customer, are closer to product needs, and no longer have to go via engineers to have an impact on the product. In a previous world, they would have maybe written down the spec and discussed it with engineers, and an engineer would have done the implementation. But increasingly, product managers, domain experts are able to affect product and be involved in product development much more directly.
Matt Turck [15:02] And then on the latter, more like the business user side of things, which, thanks for clarifying that it's not what you cover, but just out of curiosity, have you observed any kind of best practice around what you've seen customers do on prompts? Like people in whatever marketing or finance or HR, any way to guide them on how to best use those products in a way that just minimizes issues, or is that still very new territory and people have—
Raza Habib [15:21] We do have a couple of customers who use Humanloop in that way, because I think that actually the market is missing an appropriate solution. And so they kind of hack what is not built for that for a slightly different use case. But we have some customers where actually they have maybe 1,000 users who are on Humanloop and where they're saving common prompts and workflows on a per-team basis. So they say, okay, here's how you solve X task with an LLM, and there's a predefined prompt, and then they go and they open it up, they load it, they dump in their data, they run that process.
Raza Habib [15:54] Honestly, I think that it's something that companies are still figuring out, and there isn't a very good solution on the market yet. It's something that I would hope ChatGPT Enterprise just solves at some point. Things like their custom GPTs, I think, are a step in this direction of allowing people to save workflows and prompts and share them with the team. I haven't seen other good solutions yet. I'm reasonably confident that it's not well solved because of the fact that I see people kind of using and abusing Humanloop in this way.
Matt Turck [16:07] Yes. Interesting and funny. So you are able to evaluate and monitor all sorts of models, it seems, whether closed source or open source.
Raza Habib [16:14] Yeah. So we're LLM-agnostic. We sit on top of most of the major players, as well as your own models or open source as well.
Matt Turck [16:26] And is that because, from a prompt engineering perspective, it doesn't really matter how it's deployed because the way of interacting with the model is through the prompt? So the only thing that matters is what comes out as a response, or is there more?
Raza Habib [16:51] Yeah, more or less, right? The base models are somewhat interchangeable. They have different characteristics. They have slightly different performance abilities. Larger models are slower and more expensive. Smaller models are faster, but maybe less good at reasoning tasks. They go through reinforcement learning from human feedback in different ways. So the different models almost have different personalities, but fundamentally the mode of interaction is the same. The challenges around evaluation are the same. And also, many companies want to have the ability to move between them on a different use-case basis.
Raza Habib [16:59] And so it makes sense to sit on top of all of the different models.
Matt Turck [17:06] Yeah. Any lesson learned or patterns we see across different types of models in terms of—
Raza Habib [17:33] Yeah, I would say that it's still the case in terms of raw performance: GPT-4 and OpenAI's models still have a little bit of an edge over their closest competitors. Anthropic's models are also extremely strong, but amongst those closed-source APIs, OpenAI's models for tasks that require higher amounts of reasoning are still the best. And so what we see a lot of is people starting with that. It's the fastest to prototype with. It allows them to get a first version built quickly, get it to market, get feedback, validate an idea.
Raza Habib [18:03] But oftentimes the downsides of it, and OpenAI is working on this, are costs and latency. It's a bigger model. It's more expensive. It's very general. And once you do have a specific use case, maybe you don't need the full power of GPT-4. And so, once you've developed it, a very common pattern is people will build the first version with the most powerful model, and then they will try and fine-tune a smaller model for that task. And they'll often use the production data that they generated as the fine-tuning data.
Raza Habib [18:31] It might be fine-tuning GPT-3.5 Turbo. It might be fine-tuning an open-source model. The other thing that we've been seeing is just that the quality of the open-source models that are available is going up very quickly. A year ago, the gap between the best-performing open-source model and the best closed model was enormous. Now it's still big in tasks that require complex reasoning, but for many tasks, you can get away with a smaller model. And that has the advantage that if you care a lot about privacy or you care a lot about latency, you can run these things locally.
Raza Habib [18:40] And oftentimes they're much smaller.
Matt Turck [18:50] And for those open-source models, which one are the best? Is that Llama 2? Is it Falcon? Is that, like, anything you've seen?
Raza Habib [19:16] Llama 2 is still the one that I think has the best overall performance, but there are good smaller models coming out as well. So I think there's been a lot of excitement recently around the Mistral 7B model. So this 7B being 7 billion parameters. So this is roughly 10 times smaller than GPT-3, but still getting similar performance. And so that's one that's quite exciting for being both excellent and small, and also just a very impressive team. Falcon was state of the art when it came out, but I think has been superseded since then.
Matt Turck [19:27] Okay, great. I saw you had announced something called tools. Did we cover that already?
Raza Habib [19:28] Is that—
Matt Turck [19:31] So we talked about prompt engineering, we talked about fine-tuning.
Raza Habib [19:52] So we haven't actually covered tools yet. LLMs by default are text in, text out. And so they're limited in what they can do in terms of taking action in the world. And the idea of tools, or what OpenAI originally called function calling, and we sort of called tools—and now I think we won the war on the naming because they've renamed it to tools now as well. But essentially what this is, is that if an LLM wants to take an action in the world, you essentially allow it to do that by exposing a set of APIs.
Raza Habib [20:17] So, a set of different programs that the LLM can use. And you say to the model, hey, these are the tools that are available to you. So, for example, web browsing might be a tool that's available to the model. And the model then can output a request to that API, which is just a JSON string. So it's just another piece of text. And that JSON string will specify which tool it wants to use, what question it wants to ask of that tool, or how it wants to use that tool.
Raza Habib [20:47] That output is then taken, used to run the tool itself. So maybe you go and do a web search. The result of the web search is then passed back to the model, and the model then uses that output to make another decision. And so suddenly you go from a system that's only text in, text out to something that can be action-taking and that can actually take advantage of external APIs in the world or do information retrieval. So the simplest and most commonly used version of tools, I think, is just retrieval, where a model will decide to use a knowledge database to answer questions factually.
Raza Habib [21:15] That people are pushing the limits of this in very complex ways. I think I saw that you had the founder of Lindy AI on here recently. And Lindy is a great example of this, where they're making very heavy use of agents, or LLMs that are able to actually take actions in the world. And this is the way that it's typically achieved.
Matt Turck [21:23] Okay. And so this is one of the areas where you're going as well, into monitoring agent chains and to be—
Raza Habib [21:37] Yeah. So you could, by default, monitor an agent, a chain, just as well as you would a normal single LLM call with Humanloop, but you can also, in our interactive environment, be prototyping with different tools, with agents, with chat-based models that are able to take action.
Matt Turck [22:04] Okay. So circling back to some of the beginning of the discussion in terms of where you fit in, which categories, what's super interesting for people who are spending time in the space, not as technical experts, but as people trying to make sense of how it all works together, is this explosive pace of innovation and all those new different companies and categories. And each company seems to have its own evolution. And as a result, it's a little blurry who does what. So any thought there would be very helpful.
Matt Turck [22:21] In particular, frameworks like LangChain, where does that fit? Are you a competitor? Are you a partner? And then ChatGPT Enterprise, or what OpenAI does in terms of getting further into the enterprise, same thing. Is that a competitor, friend, foe?
Raza Habib [22:43] Yeah. So I think at base, you have the foundation model providers, right? So here you would have OpenAI, Cohere, Anthropic, Mistral, whoever else it might be, the open-source models, Llama. And then at the other end, the other end of this spectrum maybe, are the applications, which are the end-user-facing applications. And Humanloop is kind of a layer that sits in the middle. So we're model-agnostic. We have close partnerships, and we'd definitely be friends with all of the foundation model providers.
Raza Habib [23:08] We're keen to basically help their customers get to value. LangSmith and LangChain. LangChain is like an orchestration library. So this is basically just a set of utility tools for people who are writing the code around an AI application. And it has a whole bunch of helper functions built in that help them get started more quickly. So that wouldn't be sort of directly competitive. Their LangSmith product probably has a little bit of overlap with us, but it has some overlap, but it's fundamentally, I think, focused more around monitoring chains and agents.
Raza Habib [23:34] And then where I think we are focused is for enterprises who are building LLM apps, where typically there's a lot more collaboration required. This becomes a team sport. And also where the need to have guardrails and evaluation starts to become much more significant because they're operating at a higher level of stakes. And so I think one of the differentiators for us right from the start is that we've been focused on larger companies. As soon as you're building at the scale of a company, like, I'm trying to think of the ones I'm allowed to mention, but we have some very large companies who use us, who are financially regulated, et cetera.
Raza Habib [24:12] For them, being able to have very robust evaluation is just absolutely critical. And also they tend to have collaboration needs that are not seen at the smaller end of things. So startups and 30-, 40-person companies, or people who are hacking on the weekend, tend to have less needs around collaboration. So that's changed the nature of the products that we've built.
Matt Turck [24:33] Speaking of customers, you have some impressive and very recognizable logos on the website. So congratulations on all of that, the early traction. What are you learning in terms of who's a good customer and, yeah, what level of sophistication, what do they need to have before they come to you?
Raza Habib [24:52] Yeah, so the sweet spot for us is a kind of mid-sized enterprise company. So if they're small startups, maybe fewer than 200 people, then typically they don't care as much about evaluation. They're willing to take more risk in production and let things go wrong a little bit. And also collaboration is less of an issue for them because they're such a small team. They can all sit in the same room together and they can figure it out.
Raza Habib [25:19] On the very large end, once companies are hundreds of thousands of people, we tend to be able to solve problems for them, but the procurement cycles get very long. And so it's not where we end up focusing. And then within those companies at this enterprise level, the thing that's a surefire sign that they're going to be a good customer for Humanloop is they've reached the point where they've started actually building. There's a lot of people who have interest in GenAI. There's a lot of talk, a lot of excitement, but a smaller subset who have actually put their hands on the keyboard and written some code.
Raza Habib [25:49] And within that subset, the ones who are best are the ones who have started trying to build some kind of internal tools themselves. So almost everyone who's building an LLM application will hit these problems. They'll need evaluation, they'll need logging, they'll need monitoring, they'll need prompt versioning and management. It's not a question of if they need it; it's really a question of, do they build it themselves? Do they buy it? Which tool do they buy? But they will need it.
Raza Habib [26:10] And so for us, usually the sweet spot is a head of product from the AI team or someone senior on the engineering side books a demo with me and they say, "Hey, we spent the last two months trying to build this. We're realizing it's more complicated than we thought. Actually, we would like to just buy a solution. Can we get a demo?" And so it's the people who are problem-aware and have started building. That's really when we have the quickest traction.
Raza Habib [26:13] Yep.
Matt Turck [26:20] And they, at this stage, they typically find you in terms of go-to-market motion, mostly inbound, bottoms-up?
Raza Habib [26:33] Like 90% inbound. I would say inbound, but sales-led. So not bottoms-up. The motion isn't typically that an individual developer is playing with it, and then they buy a license and then a few more people join.
Matt Turck [26:33] Yeah.
Raza Habib [26:46] It's usually that someone sees it within the company, they come, they do a demo, and then they typically buy it for their department because they know that they're going to need a certain number of seats and usage and volume. And it makes sense for there to be a unified approach. I think what a lot of companies are worried about is there's lots of different teams just spinning up lots of things in very ad hoc ways and building their own tooling and using spreadsheets.
Raza Habib [27:14] They don't have oversight. They don't have version control. They don't have role-based access. They're worried about what data people might be putting into the models. So they actually are looking to centralize it and have more standardized processes where teams can learn from each other. And so Humanloop also becomes that central place where they're storing all their prompts, but they're also actually sharing learnings across teams.
Matt Turck [27:33] Interesting. Are you finding yourself doing a lot of evangelizing and education? And I mean, obviously a software company, some level of services, not in terms of your strategy as a company, but more as a testament to where the market is and people need a lot of handholding and help.
Raza Habib [27:54] We do a certain amount of enablement, but I don't see it as that different from other enterprise software products, right? There's a certain amount of training and help that's needed to onboard a team of people to a new product, but there are enough people now in the market who are problem-aware that actually we don't need it so much. If you'd asked me the same question six months ago, it was a very different story. But I think we're actually reaching a level of maturity in the market where people know what they want.
Raza Habib [28:09] They've tried to build something. The problems are becoming clear. There's an emerging set of tools that are needed. And typically they come to us knowing more or less what they're looking for.
Matt Turck [28:34] Well, interesting. Anecdotally, are you seeing that evolution into production? I think we've all seen this: a lot of talk, a lot of excitement, some consultants out there making a lot of money guiding companies towards potential adoption. But are you seeing some exciting use cases in production just yet, or is that in the making? What's your sense of the reality of the market right now in the enterprise?
Raza Habib [28:53] So we're starting to. It's been a little bit slower to get things to production because there's more security concerns and legal concerns and just more hoops to jump through. But we're beginning to see it now. A couple of—I can't name names for some of these—but larger, financially regulated companies are starting to use these for internal operations, where they have huge teams of people who are doing a lot of document processing, and they're able to dramatically accelerate some of those workflows using LLMs.
Raza Habib [29:04] Okay.
Matt Turck [29:38] Very, very cool. So maybe taking a step back from Humanloop, I'd be curious about your thoughts as a deep industry practitioner on some of the key debates in the industry right now. Certainly, as we're recording this, there seems to be a tweet every second on open-source versus closed-source models, and people that feel very strongly about preserving a very free, open ecosystem and others that are more concerned about security risk. Any thoughts on that debate? Where do you land? Just curious.
Raza Habib [30:00] My natural inclination is always to be in favor of open source, right? Just as a default knee-jerk reaction, we've got so much benefit from open source in general. The software world is built on top of it, so I always start from a position of optimism about open source. I understand, though, some of the reasons why people have safety concerns around larger models. There is an opportunity for misuse. And I've seen people like Yann LeCun say, "Oh, but we already have search engines."
Raza Habib [30:24] So people can look up with a search engine how to build a bomb or how to build a bioweapon. Why do LLMs make it worse? But I think that really does downplay how much better they are at synthesizing information and explaining steps to you. And there is a dramatic reduction in how hard it is to do certain forms of misuse. And that's before we get to the safety concerns about AI that actually is misaligned.
Raza Habib [30:51] But to me, the question is, it's very difficult to know how to solve this or when to put restrictions in place. When GPT-2 came out, they didn't release GPT-2 for fear of misuse concerns. And then GPT-3, and now we're on GPT-4. It would have been really sad for the world, I think, if at the point of GPT-2, we had decided, "Hey, this is too dangerous. No one can have access," because we would have missed out on all of the past two or three years of incredible progress.
Raza Habib [31:20] And knowing when that moment will be or how to draw it feels to me like it's almost never going to be clear that you can say, "Okay, actually this is too dangerous now." And so I'm generally in favor of finding ways to regulate end use cases and make the malicious use of the models regulated, and then punish that very strongly and make people responsible for the end outcomes, rather than banning the underlying technology. Though I keep a door open to the possibility that we will reach a capabilities threshold where actually that's no longer a tenable position to hold, where the models are so powerful that you really don't want this to be in the hands of any random citizen in the world.
Raza Habib [31:51] I've heard arguments from people like Jeremy Howard who worry about centralization of power. So they say, if the models are very dangerous, one of the strongest counterarguments is that we don't want a small number of private companies or governments to be the only ones who have access to this. He cites the Enlightenment arguments of, during the Enlightenment, people said the ability to read is something that should be more widely distributed, and that gives people power and is actually better for society.
Raza Habib [32:25] I think there are flaws in those arguments when the models become particularly powerful. When you would not want the average citizen to have the ability to create chemical weapons at home, you would not want the average citizen to be able to—and I'm not saying that LLMs allow them to do these things—but I'm just saying that the argument fails in a particular way. There is a capabilities threshold above which it is correct that we would want to restrict access.
Raza Habib [32:44] And to me, the challenge is knowing when that threshold is. And I would like to be very cautious in restricting open-source usage ahead of that, because the evidence of the past few years would suggest that it's easy to do it too early. And I think we're still figuring that out.
Matt Turck [32:59] Very, very cool. All right. Well, it's been a really wonderful discussion. Maybe to close, what's next for you guys? The next six months, year, two years? What are you perhaps building? Where do you want to be? Where do you see things going?
Raza Habib [33:22] Yeah. So one is doubling down on the areas of strength we have today. So continuing to try and be the best player for evaluation, for helping people when they're doing all of their collaborative prompt development and engineering, versioning, managing. But then once we've really nailed that down and have become the leader in that space, I think the directions of travel—one is more proactivity. So today, a lot of these monitoring and observability tools are passive. They do observation, they help you find bugs, help you correct them.
Raza Habib [33:42] But actually, we're building AI tools. They can do a lot more than that. They can actually start making suggestions: hey, here's how you can reduce costs, or here's how you might improve a particular version of a model. Proactivity will be a big part of it. I also think increasingly, the prize is trying to get agents to work reliably. And as the models improve, I think that's still some way away, but better supporting people who are building agents will be something that becomes more and more important over time.
Raza Habib [34:14] In the short term, it's really nailing our core competencies of being best in class at evaluation and having the best collaborative environment for prompt engineering, rolling that out to more and more companies, because I think we are still at the early stage of the adoption curve there. And then I think within the next six months, moving on to more proactive versions of this, where you're not just passively monitoring, but are actually helping make things better in an active way.
Matt Turck [34:21] Great. Where can people find you online, learn more about you, learn more about the company, follow your thoughts?
Raza Habib [34:37] Humanloop.com is the website. There's docs there. You can sign up and try the product for free. There's case studies. That's probably the first place I would go to find out information. We're also active on Twitter and LinkedIn. I'm on Twitter @razrazqol. And other than that, I've also done a few of these other podcasts. So if people are interested in finding out more about me personally or about the company history personally, there's a good set of content out there.
Raza Habib [34:44] Great.
Matt Turck [34:48] Raza, thank you so much for doing this.
Raza Habib [35:16] Thanks for joining us for The MAD Podcast. We're back here every Wednesday with new conversations with leaders in the machine learning, AI, and data space. And if you like this show, you can also find a video recording of not only this episode, but many, many more over on the Data Driven NYC YouTube channel. Thanks again, and catch you next week.