Ex‑DeepMind Researcher Misha Laskin on Enterprise Super‑Intelligence | Reflection AI
The MAD Podcast with Matt Turck · with Misha Laskin, Co-founder and CEO, Reflection AI
Misha Laskin is the Co-founder and CEO at Reflection AI. We cover why organizational superintelligence should function as an oracle, why code agents without context become L9 engineers with amnesia, and how Asimov combines code, chats, docs, and project tools into team-wide memory.
Chapters
- 1:42 — Reflection AI: Company Origins and Mission
- 4:14 — Making Superintelligence Concrete
- 6:04 — Superintelligence vs. AGI: Why the Goalposts Moved
- 7:55 — Organizational Superintelligence as an Oracle
- 12:05 — Coding as the Shortcut: Hands, Legs & Brain for AI
- 16:00 — Building the Context Engine
- 20:55 — Capturing Tribal Knowledge in Organizations
- 26:31 — Introducing Asimov: A Deep Code Research Agent
- 28:44 — Team-Wide Memory: Preserving Institutional Knowledge
- 33:07 — Multi-Agent Design for Deep Code Understanding
- 34:48 — Data Retrieval and Integration in Asimov
- 38:13 — Enterprise-Ready: VPC and On-Prem Deployments
- 39:41 — Reinforcement Learning in Asimov's Development
- 41:04 — Misha's Journey: From Physics to AI
- 42:06 — Growing Up in a Science-Driven Desert Town
- 53:03 — Building General Agents at DeepMind
- 56:57 — Founding Reflection AI After DeepMind
- 58:54 — Product-Driven Superintelligence: Why It Matters
- 1:02:22 — The State of Autonomous Coding Agents
- 1:04:26 — What's Next for Reflection AI
Transcript
Reflection AI: Company Origins and Mission
Matt Turck [1:47] Hey, Misha, welcome to The MAD Podcast. Thanks for being here in person today.
Misha Laskin [1:49] Yeah, thanks, Matt. Thanks so much for having me.
Matt Turck [2:10] You are the co-founder and CEO of Reflection AI, which is a super exciting startup that, up until recently, was pretty much in stealth. You sort of came out of stealth a few weeks ago, and you have some exciting product announcements that we're actually going to discuss today. Before we do that, tell us about the company itself.
Misha Laskin [2:37] We're a research and product company in AI that was founded by myself and my co-founder, Ioannis. We are both longtime researchers at DeepMind, and generally the makeup of the team is a lot of folks who did large language model training and reinforcement learning, building reinforcement learning systems within companies like DeepMind, OpenAI, and Anthropic. We came together to build this company that co-designs research and product to build superintelligence. I think that superintelligence can sound like a very abstract thing, and I think what makes doing this at a startup or a smaller place special is that you can be very concrete in what that means to you.
Misha Laskin [3:12] Because I think that generally, when people think about superintelligence, they might think of something that becomes really good at math Olympiads or coding Olympiads, competitive coding, things like this. But these things to us don't really—I mean, they're impressive, but it's unclear how they're useful. Like, what is the superintelligence doing that is useful for kind of an end user? But if you co-design a product and the research alongside it, you can be much more opinionated about where it is that you want to go.
Misha Laskin [3:48] And after talking to a lot of organizations and enterprise customers after we started the company, we realized that probably the form factor of an organizational superintelligence, like the thing that's really going to help organizations get a lot of stuff done, is probably going to be something like an oracle. Like an oracle that understands the entire organization extremely deeply, that can answer questions at the level of, let's say, the senior-most person on that team, right? It kind of has that, like, a principal-level engineer or the senior-most sales leader who has full context on the company.
Making Superintelligence Concrete
Misha Laskin [4:15] And it'll probably be superintelligent also in the sense that it's doing all those things at once, right? As opposed to having these distinct different functions. And so that is what we think an organizational superintelligence is going to look like, that it's going to be a system that really deeply comprehends the organization.
Matt Turck [4:47] It's a great term, organizational superintelligence, as opposed to the super abstract concept of superintelligence. Just to stay for a second on the concept itself: superintelligence. Does the industry actually agree on what that means versus AGI? So we're recording this. Obviously, the big news is our friend Zuckerberg showering top researchers with extraordinary amounts of money and poaching them from various organizations to create a superintelligence lab, and then Safe Superintelligence as a startup. Do people agree on what that means, or is that just more like something that directionally feels very impressive, but not everybody exactly knows what that means?
Misha Laskin [5:18] I would actually say that superintelligence in these contexts of the large lab context is actually just being used synonymously with what AGI used to be used for. There's a scene in Indiana Jones and the Raiders of the Lost Ark, the first scene where he goes and tries to steal a sculpture, and he kind of takes a sculpture and replaces it with a sandbag and hopes that nothing changes. And I kind of think one of those—
Matt Turck [5:23] That scene does not end well, if I remember correctly. It doesn't end well, but the analogy stops here.
Misha Laskin [5:46] Yeah, the analogy stops here unless we really figure out safety. But I think that what happened is that, by many accounts, many people would argue that AGI hasn't been achieved, but now some people might argue that AGI has. And so it's unclear that this is a binary event that hasn't been achieved yet, even though that was the case before. And so I think the goalposts were just moved. It's like, oh, what I meant by AGI was actually superintelligence.
Matt Turck [5:47] Sure.
Misha Laskin [5:58] So I think that is to say that it's as vaguely defined as AGI was before. And I don't think there was any agreement on what we actually meant by AGI.
Superintelligence vs. AGI: Why the Goalposts Moved
Matt Turck [6:04] Right. It does sound cool, though. 2025 superintelligence is 2024 AGI.
Misha Laskin [6:05] Okay. All right.
Matt Turck [6:23] Yeah. So coming back, I really like that term of organizational superintelligence. It does sound more tractable. And your analogy of what a senior person would know, especially, I assume, a senior person that actually does the work. Right. Because typically, the most senior people in an organization are actually disconnected from the reality of what happens in the trenches.
Misha Laskin [6:50] Yeah, I think it's more the person who—on every team, there's a go-to seasoned, experienced person who's very much in the weeds and understands everything. Getting systems first that are like that, and then they can hold a lot more context in their heads and sort of be even better than those critical members of the team. I think that's kind of like an oracle for an organization. That's what a superintelligence looks like. Because once you deeply understand the problems that you need to solve in any discipline in enterprise, acting upon them to actually go and solve it is the easy part.
Misha Laskin [7:10] And so we're really focused on coding to start for a number of reasons. We kind of think you solve this problem, it sort of solves the more general problem.
Matt Turck [7:11] Yeah. What are those reasons?
Misha Laskin [7:22] One reason is that the way today we think of coding is just for software engineers. But when you build a coding model, you've just trained a language model that can interact with a piece of software through code. And so the way these language models are going to interact with any piece of software, not just software engineering software like Salesforce and other CRMs and creative tools and so forth, the majority of those interactions are going to be through function calls, APIs, so through code.
Organizational Superintelligence as an Oracle
Misha Laskin [7:56] So today we think of coding models as something that is just for software engineers, but that's not true. If you build that system, it can actually interface with any piece of software. It's sort of what you've built is kind of the hands and legs of a digital AI, right? In the same way that a humanoid robot can do anything with the same hands and legs that humans do.
Matt Turck [8:04] So would coding be the brain, and then everything else becomes the legs and the hands that are connected to the brain? Or is that the wrong analogy?
Misha Laskin [8:24] Yes, you can kind of think of the model as the brain, and then what people call scaffolding, agent scaffolding, that's sort of the affordances, right? The things that you can actually do. All these affordances for software, at least for digital intelligence, are going to be primarily through code. That is, the other option is teaching a model how to drag a mouse around, and it's called computer control. Some of that will happen, but I just don't think that that's going to be the majority way in which a language model interacts with software.
Misha Laskin [8:35] So if you solve coding, you've just solved how a language model should interact with software?
Matt Turck [8:48] And also, it's a more tractable problem. Is that fair? Because code is more structured and code is closer to grammar, lends itself well to the type of work LLMs do. Or is that—
Misha Laskin [9:14] No, I think that's very much correct. It's more that code is kind of intuitive to LLMs because they were trained on the internet initially. And so they kind of have almost a muscle memory for code. It's kind of like humans have evolved to have a spatial, geospatial muscle memory. And so we learn very quickly how to use our hands and so forth as toddlers. Code is actually very unintuitive to us, right? That's why the GUI was invented, because the first way to interface with computers was through code.
Misha Laskin [9:46] Very unintuitive for humans. So you invent the GUI. For language models, it's the opposite. They've never seen any GUI data really on the internet. There's no mouse movement data out there. So the thing that is native to them is the exact opposite of what's native to us. So if you build things that are really good at coding, you're sort of just amplifying the existing base knowledge of the LLM. It's just intuitive to a language model. So that's why coding.
Misha Laskin [10:10] But then within coding, you can kind of think about what it would take to create a superintelligent coding agent. You basically need a system that can generate code really well, and you need a system that comprehends large codebases and all the knowledge around them, like everything that's in all the docs and the stuff that's kind of tribal knowledge that's in engineers' heads. And I'd say every coding tool today is focused on the code generation piece in an autonomous, semi-autonomous, in-your-IDE, in-your-terminal way.
Misha Laskin [10:39] But it's all around code generation. And as a result, because companies today, I don't think, are really focused in earnest on the understanding piece, what we're building towards, I'd kind of call as almost today, maybe we have AI engineers that are at the level of an L4 engineer, maybe next year L5, L6, going to L9. I think the path there is clear. And what we're going to get to if we don't solve the comprehension piece is basically L9 engineers with amnesia.
Misha Laskin [10:51] So an L9 engineer, if they came into your company but had amnesia, they wouldn't be really useful.
Matt Turck [10:52] Good news and bad news.
Misha Laskin [11:11] Right. It's kind of, if you tell me exactly what to do, I'll go do it really well. But otherwise, I know nothing about your codebase, I know nothing about your organization, and I'm not going to gather any of that either. The thing to solve now, I think, is building what we kind of think of as the context engine. Like, what's the thing that's kind of the brain that your coding agents can query in order to have all the context that they need to go do meaningful work?
Misha Laskin [11:45] Not just junior engineering work, but the stuff that engineers spend really most of their time on, like critical infrastructure bugs, dealing with legacy code, things that if you actually look over the shoulder of an engineer at any large organization, or even a startup with a really sizable codebase, and you look at what they do, you'll find that a minority of their time is actually spent coding. Like 70% of their time, they're digging basically through information, trying to understand stuff, asking other engineers questions.
Coding as the Shortcut: Hands, Legs & Brain for AI
Misha Laskin [12:06] And you kind of want to build a superintelligence that has that same DNA. What it's spending most of its time on is collecting information, and then obviously, once it has it, it knows how to act on it.
Matt Turck [12:20] How is that different or similar to basically doing a combination of coding and RAG, and RAG to bring in the context of the enterprise? Is that the general direction?
Misha Laskin [12:37] Maybe it's almost like if RAG worked. So the purpose of RAG is to get all the information that your coding agent needs in order to do some work. But it's a very primitive form of doing this, and by and large fails. It's kind of—
Matt Turck [12:43] I mean, why is it primitive? Because it's just search-oriented. It's just a query, or it's more—
Misha Laskin [13:07] So there are a few things about it, but the first thing is that it's kind of what we call sparse, because the stuff that it grabs from a large codebase has a lot of false negatives and false positives. And it typically only does it once, right? It'll grab it, and then that's all you have. And most likely, for any meaningful query, it will not have given you the information that you need to actually go do the task. So RAG agents actually are pretty weak.
Misha Laskin [13:33] What's happening now is that there's kind of a new kind of retrieval that I think people are calling agentic search, which is more what Claude Code does. And it's kind of an agent that uses the same kind of command line that a human does and goes and kind of looks for files and uses the same commands that an engineer does to go look for files and search for keywords and grab stuff. And it does this agentically. So as opposed to it just being a one-step process where it just pulls a bunch of embeddings in, it will use a file search tool, and then it'll look at the stuff that it did and think about it, and then use another tool and so forth.
Matt Turck [13:49] And will remember it? Is there a concept of memory built in?
Misha Laskin [14:16] It does store the stuff in a context. But the way I would think about it is that imagine you are dropped into a large, dark jungle, and all you have is, like, a tiny flashlight. That's basically what agentic search is today, where you're going and you're kind of exploring, and you have this tiny flashlight, and you have to then remember everything in your head. And obviously, that's a first form of comprehension, but it's also a pretty weak form of comprehension. It doesn't scale to large jungles.
Misha Laskin [14:24] Like, if you have a little tiny jungle in your backyard, then you might be able to navigate it.
Matt Turck [14:25] A bonsai jungle.
Misha Laskin [14:48] A bonsai jungle. Yeah. So that's kind of how I liken the existing state-of-the-art agentic search, whereas the kinds of systems that you need to build are ones that basically expand the kind of lens, like the aperture of your beam, so you're seeing more, are smarter about where they go and look for stuff, and how they remember it. These are basically fundamental research problems, like memory. Long-context reasoning is what people call it, but I'd say it's just expanding the aperture of your beam, figuring out how to actually store it, and also what sources of information should you be pulling on other than just your code.
Misha Laskin [15:27] Because information for organizations lives in their chats and their project management tools, and a lot of it is in their team's just brains, right? So oftentimes, when that senior engineer leaves who understands some legacy piece of code, all of a sudden that becomes known as a haunted graveyard because no one else understands it. And so how do you build products that capture that knowledge that's in people's heads? That's why I kind of think that you can't—it's so hard to think about superintelligence in the abstract because half of it is a product problem.
Building the Context Engine
Misha Laskin [16:02] The comprehension problem is sort of, I would say, deep. And what people used to say, like, when a problem was AGI-complete, it just meant that the problem has so much depth to it that if you solve it, you get AGI. I would say this is a superintelligence-complete problem. Like, if you really solve this oracle for organizations, just for coding, you've basically built all the capabilities you need to have a superintelligence.
Matt Turck [16:34] Because you can generalize from there. It's an idea that I've certainly seen make the rounds, by the way, that coding is a path to superintelligence. I think a variation of that idea has been that it's intelligence building intelligence, doing it automatically, which I think feels like a different avenue. But what you're saying here, or maybe not, but I think what you're saying here is more that it's such a complex problem that if you solve that problem, then you're there. I mean, just to play back what you just said.
Misha Laskin [17:03] Yeah, I think that the notion of, in the sense in which coding is ASI-complete, is that you can use an intelligent coding system to build another intelligent coding system. If people like to do research in the abstract, that sounds awesome, because it's kind of like, I don't have to think about the actual problem that's being solved. If I just build an intelligence that builds an intelligence, it will figure it out somehow. I don't think that's how it works. I think coding intelligence that builds better coding intelligence will make algorithms more efficient, basically.
Misha Laskin [17:38] And maybe it's so intelligent that it can even start making all the product decisions for you in terms of what questions you should be asking users and what features you should be building. And to me, it actually seems like a much larger, more practical problem. There's almost no—I think research without co-designing product with it is sort of a meaningless pursuit now. That was only meaningful when the ingredients for how to build artificial general intelligence, or ASI, were not known. I think now they're known.
Matt Turck [17:46] So you think that we have everything? And maybe list what everything is in terms of components.
Misha Laskin [18:05] Yeah, I think we have everything, even though the capabilities will continue being tightened. But the sort of series of breakthroughs that need to happen, many of them have happened. There are still some, but I would say that the ones that come to mind, and there, I'd say, are four of them. The first ones were just making deep neural networks work. So that was just the ImageNet moment in 2012, when you can have a deep neural network that's classifying images at human and then superhuman level.
Misha Laskin [18:39] I think that was the first breakthrough. Second breakthrough was reinforcement learning, even though I think it was not clear maybe then, but now it's clear again. So reinforcement learning basically told us, how do we take a system, and if we have a reward, make it superintelligent? We actually—superintelligence has been built a few times now. AlphaGo was superintelligent. AlphaStar was—if you put more compute into it, it would have gotten superintelligent. So we've known how to build narrow superintelligence for a while.
Misha Laskin [18:45] So deep neural networks, reinforcement learning.
Matt Turck [18:50] So 2012, AlphaGo was 2016 or '19? Yeah, 2016.
Misha Laskin [18:54] Then scaling up transformers, the GPT series of work.
Matt Turck [18:55] 2017, '18.
Misha Laskin [19:19] Kind of '18, but then really, I think, started shining in the early 2020s with GPT-3. And that was both an architectural innovation and a data innovation. It was an architecture that could consume data on the internet. So it's kind of both hand in hand. Transformers are not just architecture; it was also data. And then I would say the last ones are RLHF, which was, okay, we know how to align models in the basic way. We know how to do basic reinforcement learning on top of language models.
Misha Laskin [19:41] And then reinforcement learning coming back again with the reasoning models. So if I was to summarize, I would say neural networks, deep neural networks, transformers and internet-scale data, and reinforcement learning in kind of its various incarnations. These are the ingredients that are required to build a superintelligence.
Matt Turck [19:49] And as a return of reinforcement learning at the end that you described, does that obviate the need for RLHF?
Misha Laskin [20:16] Was that a temporary solution, or are those parallel? I think they're parallel capabilities because they do different things. RLHF was more—the design there was more to align a language model, like a pre-trained language model. If you play with one of these base models before they're aligned, they're really useless. They feel like stochastic parrots. They don't follow instructions. They're extremely kind of high entropy. And the fact that you could just align them, tweak them to align them with something that is human-consumable, I think was a pretty big breakthrough.
Misha Laskin [20:52] And that was RLHF. Whereas RL with reasoning is more RL to drive intelligence capabilities, which is make it really good at coding, make it really good at math, make it really good at whatever target domain that you have rewards for. And you use these in tandem, where RL in the reasoning phase sort of really expands on a capability, and then RLHF aligns it to be kind of human-consumable. But they're the same thing. In fact, I would say it's just the same thing.
Capturing Tribal Knowledge in Organizations
Misha Laskin [20:56] The machinery is the same.
Matt Turck [21:14] So we have all the components for AGI, SAI. Your goal is to create this in an organizational enterprise context. And I guess let's get into the big news of this week, which is that you're launching your first product. Tell us everything about it.
Misha Laskin [21:41] Yeah, we're launching our first product, kind of the first milestone on the path to superintelligence. It's called Asimov, like the science fiction writer who had some thoughts on the subject. And Asimov is the best-in-class code research agent built for organizations. So it's different than a coding agent, which is a thing that will go and write code for you, but which today we feel is pretty context-poor. So if the task is well-scoped, coding agents today will do it pretty well for you.
Misha Laskin [22:16] If you are spending most of your time trying to understand some hairy problem in your engineering, why some infrastructure bug is happening, or really doing what engineers spend most of their time doing, which is this kind of work, there's no tool today that really unblocks them in that way. And so I'd say 70% of doing software engineering is actually a code research job rather than a code monkey job. And we're building a product that addresses that problem. So the majority of time that engineers spend unpacking problems and trying to understand why something's happening is oftentimes bottlenecked because they have to ask someone a question, and it takes a few hours for that person to get back because they're busy.
Misha Laskin [22:38] We've built a product that helps overcome these kinds of problems.
Matt Turck [22:42] So just to play it back, it's like deep research for code.
Misha Laskin [23:03] It's very similar in concept to deep research for code. A thing that will go explore your codebase and other sources of knowledge. It might take a bit longer than the traditional snappy ask-kind-of products, but it'll come back with much better answers. Things that make it different are that, one, engineering knowledge does not just live in a codebase. It lives in all sorts of other surface areas and other software products, like project management tools, chats, documentation, things like this.
Misha Laskin [23:37] And Asimov pulls from those things. So it's not just your codebase; it's aggregating all these sorts of information. It has this new concept that I don't think has been introduced to date that we call team-wide memories. So today, products have individual memories, which remember the personal preferences of the developer, but they don't have a team-wide organizational understanding. Suppose you have a senior software engineer that understands some microservice A there. You are not that engineer, and now you need to interact with their microservice, and you don't have the context they do.
Misha Laskin [23:51] So now with Asimov, it kind of organically captures that information as it happens in chats, but also engineers can just teach it directly.
Matt Turck [24:02] That's super cool. And then that addresses the problem you mentioned, I think at some point earlier, which is like, if somebody leaves, then the institutional memory goes with them. So you have a permanent organizational memory.
Misha Laskin [24:22] Exactly. It's kind of building a permanent organizational system of record for your engineering knowledge to start. A side note on that is something that's been interesting is that some of the most excited users have been these senior staff-level engineers who are fielding people's questions all the time. And so what we've seen is that we'll go to an organization, and the first three weeks there'll be four senior staff-level engineers who are just populating its knowledge. It's kind of the first time I've actually ever seen engineers excited about documentation, effectively.
Misha Laskin [25:06] That's kind of a unique thing. And then the final thing is really around agent design that enables you to, what I was calling earlier, increase the aperture of your beam so that it's looking at much larger codebases than agents were able to before. This is still, in that sense, a work in progress, in that long-context reasoning is just a big fundamental problem. And I think there's a lot of work to be done there. But I think this new agent design is a step in that direction.
Matt Turck [25:15] Directionally, why and how are they able to do that?
Misha Laskin [25:37] This is not unique to us. I think when I look at how state-of-the-art agents are being designed, they're being designed very much as these multi-agent systems. Agent design is effectively a big reasoning agent, and I think that's pretty standard, but it dispatches small long-context reasoning agents to go search for different relevant chunks of information in the code. And so I think that it's this kind of decoupling of a big reasoning agent, maybe with a smaller context, with a bunch of these little retriever scout agents.
Misha Laskin [26:05] It's a different design than what's happened to date, though. I'm sure that other companies will converge on it as well. You really want to design your agents for the problems that you're trying to solve. So if you want a really snappy agent that's going to answer things immediately, then this is probably not the best design for it, right? You might want something that does RAG, which is very fast, or a very basic search agent, kind of like what Claude Code or Cursor might do, which is something that just uses the terminal and the file system there and uses the same commands as an engineer.
Introducing Asimov: A Deep Code Research Agent
Misha Laskin [26:33] That's a lot snappier than sending a lot of these retrievers out. So I think what we'll start seeing is this product co-evolving with the problems that you're solving, and there will be different agent designs for different problems.
Matt Turck [26:46] And is a sort of snappy one-shot agent necessarily a bad thing if you direct them at small problems and then you try to put them together? Or does the sort of individual little agent need to be smarter?
Misha Laskin [27:05] No, I think that it's a really interesting question. And this is kind of one of the things that's interesting in research now, as opposed to before, is that before it used to be just around models, but now a lot of the research is in your agent design. And so what you ask is kind of an open question. There could be a hybrid system that kind of routes some queries to one agentic system and then routes other queries to another agentic system.
Misha Laskin [27:46] It could be that you figured out some elegant, simpler kind of multi-agent system that can do both things. It's kind of an open research question. It's very exciting. And in some sense, it parallels a lot of the unspoken research that was happening at DeepMind and probably at OpenAI as well during pre-language models. For example, projects like OpenAI's Dota 5 or DeepMind's AlphaStar project, which trained these expert-level agents to play pretty complex video games like StarCraft and Dota. A big question there that I don't think that many people appreciated is, how do you design the environment for your neural network to actually dispatch actions?
Misha Laskin [28:23] Right. So the most simple thing you can think of is, well, it learns to use a keyboard and mouse like a human does. That turned out not to work. So when you read the AlphaStar paper, you see that they actually figured out this particular way of factoring out the actions in order to make neural networks play StarCraft. And what that actually meant is that that was agent design. So now agent design is back in terms of designing scaffolding. Back then it used to be called environment design, but it's the same thing.
Team-Wide Memory: Preserving Institutional Knowledge
Misha Laskin [28:44] And a lot of the project in these big projects was not even on training the models. It was figuring out how the agents should actually interface with the environment that you're training it in. And where do you get your data from? What's your data? How do you collect it?
Matt Turck [29:07] So on that point, you had mentioned the three things, like multi-agent design. I think the first point was what sources of data it accesses. So is there, I think, to what you just said, some difference in how those agents access data versus RAG, which is pretty much like straight-up search? I guess as a side question or related question, do those data sources need to have a special protocol to lend themselves to agents, or like MCP-style kind of infra, so that the superintelligent agents that go around and sort of grab information everywhere can interact with them?
Misha Laskin [29:44] It's a really good question. So there are kind of a couple of interesting things to unpack there. The first thing I guess I would say is that, depending on the problem you want to solve, different types of search all fall into basically search. And maybe I would say that if you want really fast, but it doesn't really matter how accurate it is, or it just needs to be some kind of ballpark accurate, but really fast.
Misha Laskin [30:14] So, kind of the weakest form of search, RAG is great. Then there's this more agentic search that we spoke about, where the agent uses tools similar to the ones available to humans on a computer. It's kind of in between the spectrum of fast and slow, so it's slower than RAG, but it gives you better answers. And then even slower is what I'd call neural retrieval, which is you have a really long-context model and you ask that long-context model to retrieve stuff for you.
Misha Laskin [30:49] You feed it everything you can, maybe use multiple of them if everything doesn't fit in one. And then you ask that model to look at what you put in its context and retrieve the relevant stuff for you. That's kind of called neural retrieval, and that is going to take the longest. It's not guaranteed to be the best, but you can train it to be really good. And so that's kind of the spectrum of search capabilities as I see them today. The difference between something like this and MCP as it pertains to interacting with different sources of knowledge is that MCP is kind of like, it's stateless.
Misha Laskin [31:21] It's just a way for you to interact with another piece of software. But what you need to do here is you actually need to collate and index knowledge, right? You need to take data from software, store it somewhere, and make it searchable for an agent. That ends up being a bit of, I would say, a blind spot. There's a question of why hasn't this been done? One of the reasons is that, just from a—forget intelligence—just from a business model perspective, it falls into a bit of a blind spot for existing coding tools, which were meant to be served as SaaS offerings for a broad consumer base.
Misha Laskin [32:01] But an enterprise, like, this is key IP. It will not want this leaving into SaaS. And so you have to kind of rebuild your entire business to basically deploy the stuff on that enterprise's resources. And so most companies have been basically staying away from this problem of indexing and integrations at this level of depth because it requires them to basically change their entire go-to-market and business model. But it's also why it's really easy to switch around between various different coding providers today because they don't really integrate deeply.
Misha Laskin [32:26] And so you can try Cursor today, you can try Claude Code tomorrow, you can switch to Windsurf. And there's really—because they're pretty thin skins, I guess, on just the language model API—from a developer perspective, it's pretty easy to switch around and try different things.
Matt Turck [32:38] Interesting. So is a consequence of that that Asimov needs to be able to work in a sort of air-gapped context, virtual private clouds, on-prem?
Misha Laskin [33:06] We do have a SaaS offering because ultimately, if you flip something into on-prem, it has to start off as SaaS anyway. But the primary benefit to enterprises is we're not going fully on-prem today, but we are doing VPC, and that ends up being sufficient for a lot of big enterprises out there who already have their cloud infrastructure on AWS or Azure or GCP. This notion of it being deployable as VPC is extremely important to our organization. This is definitely, it's really a deal breaker.
Multi-Agent Design for Deep Code Understanding
Matt Turck [33:25] You can't even start the reinforcement learning part in a simulator. How does it manifest? I mean, you guys are super world-class RL specialists. So is that the whole idea? Like, it keeps sort of learning and getting sharper with every interaction?
Misha Laskin [33:48] Let's say before language models were useful, you kind of had to be in this world where you build the best language model and then you figure out the product. That's kind of the world we were in. And Anthropic spent a few years building a language model and then took off with Claude 3. OpenAI, GPT-2 is not really productizable. GPT-3, not really either. GPT-3.5 and 4, they were able to productize effectively. We're in a different world today where language models are pretty good, and so our strategy has been: we build this kind of multi-agent system.
Misha Laskin [34:20] Some parts we're training models for. We kind of see blind spots from third-party models, and other places the third-party model stays for today. Over time, we're going to kind of abstract all of it, but we're being a bit more strategic about which parts of the system you need to go after as a startup, because that's kind of what you have as a benefit as a startup, is that you can be a lot more focused on the problem at hand. Obviously, the downside is that you have to be a lot more strategic about the bets that you're taking.
Data Retrieval and Integration in Asimov
Misha Laskin [34:49] You don't have the resources to go and train everything all at once, so you kind of have to take it one step at a time. And so in the long term, this is going to be a system that just learns end-to-end with reinforcement learning. And in the short term, we're applying reinforcement learning to fix problems that we're seeing as kind of blind spots in the existing sets of models when we're deploying them.
Matt Turck [35:26] I love the pragmatic undertone to everything that you're saying, which is really interesting and, dare I say, somewhat refreshing. Look, there's different ways of building wonderful things in AI, but the pragmatic tone is really interesting and not that widespread. So that's the product that you just launched. Maybe take us back a little bit. I alluded to your background as being world-class in reinforcement learning. How did that all come about? What was your journey to starting all of this?
Misha Laskin [35:30] As a kid, I got pretty obsessed with physics and wanted to be a theoretical physicist.
Matt Turck [35:32] As regular kids do.
Misha Laskin [35:42] As regular kids do. Well, if you're a Russian Jewish kid dropped in the middle of nowhere America, which is what happened to me.
Matt Turck [35:50] Yeah. So let's go into that. So you were born in Russia, then immigrated to Israel as a kid.
Misha Laskin [35:56] Born in Russia, immigrated to Israel as a kid, and then immigrated to the States for the second half of my childhood.
Matt Turck [36:05] Is that a thing that top people in AI do? Because I can think of other people that were born in Russia and immigrated to Israel, then came to Canada.
Misha Laskin [36:07] Yeah, Canada is a big one.
Matt Turck [36:09] Referring to Ilya, I guess, now at SSI.
Misha Laskin [36:29] I think the Soviet Union was an extremely technically academic culture. And a lot of the young scientists, when the Soviet Union fell apart, left to Israel, Germany, America, and Canada. And when we arrived in the United States, it was actually—
Matt Turck [36:33] So how long were you in Israel for? Just as a quick layout.
Misha Laskin [36:34] I was there for eight years.
Matt Turck [36:35] Eight years, wow.
Misha Laskin [36:48] Yeah, so from one to nine, I was there. I was just born in St. Petersburg, so I have no memories, obviously, of living there, just visiting. And then, yeah, arrived in the States.
Matt Turck [36:54] So where was nowhere America that you mentioned? What city was it?
Misha Laskin [37:18] In rural Washington State. So in Washington State, many people don't know this, but it has a hard line where the west side of the state is a lush forest, and then the east side of the state is a desert. And you can see exactly where that starts. It's not a gradual transition. It's just like trees, a dense forest, just starts somewhere. And then it transitions into immediate desert. I lived on the desert side. There's a national lab there.
Misha Laskin [37:35] My parents are chemists, and they got jobs in this national lab. The culture of that town was, it was one of the sites during the Manhattan Project. It was called the Hanford Site. It's where the plutonium was enriched. So it was the sister site to Los Alamos.
Matt Turck [37:36] This is a whole vibe.
Misha Laskin [38:01] It is a vibe. Everything is themed around that event. The bowling alley is the Atomic Bowling Alley. The brewery is the Atomic Brewery. The streets are like Uranium, Mercury, Plutonium. The park there by the river—it's on the Columbia River—is called Leslie Groves Park, who was the ruthless general in charge of that project. So it's an intense town. Actually, the most intense thing is that the high school mascot, the town's called Richland, and the mascot is the Richland Bombers.
Enterprise-Ready: VPC and On-Prem Deployments
Misha Laskin [38:15] So they're B-52 bombers, and there are mushroom clouds on the basketball court, and that's it. It's an intense town.
Matt Turck [38:21] So, hence, it all makes sense now that physics would be the escape.
Misha Laskin [38:48] Yeah, I guess I was—well, I didn't even think about the town's history as physics, even though that definitely is the case. But it was more that I was learning to speak a new language, had a lot of time on my hands, and my parents had their lecture books from—they bought this Feynman kind of series of lectures. And that was around, and I just spent some time reading it and just got into it. So that was the path to physics.
Misha Laskin [39:08] Ended up going through and doing a PhD in theoretical physics and actually defected to artificial intelligence. I realized that—first, I saw AlphaGo come out. And the short of it is, I realized that I had picked an interesting science, but not the science of our time.
Matt Turck [39:09] Hmm.
Misha Laskin [39:28] That was kind of it. All the things I was learning in physics, all these kind of very interesting, great things were done basically 100 years ago, 60 to 100 years ago. And when you're studying something as a student, you don't really think too much about the timelines. You're like, "Oh, this is cool. This is so cool. I want to do this." But then, 60 years later, the field has crystallized in many ways, and it's not as dynamic as AI, where the frontier is just moving so fast.
Reinforcement Learning in Asimov's Development
Misha Laskin [39:44] And so when I saw AlphaGo come out, to me it just seemed like, oh, this is actually the science of our time, and I needed to do that.
Matt Turck [39:59] And as a quick segue on that note, it was super interesting to see the Nobel Prizes a few months ago, or maybe that was last year at this point, everything converging towards AI. Do you think AI is eating all those other scientific fields?
Misha Laskin [40:27] I think it's augmenting. And in some sense, I think that part of what happened was that AI's impact in the world had clearly become hard to ignore. But there's no Nobel Prize for computer science. The Turing Awards have gone to AI for a number of years now. And so I think that the Nobel Committee, I'd imagine, felt like it needed to somehow shoehorn this. And so obviously there were very impactful AI breakthroughs with AlphaFold that resulted in a prize.
Misha's Journey: From Physics to AI
Misha Laskin [41:07] But what's interesting is that the Physics Nobel Prize was given to something that has not really had that much impact in physics. But I still buy it because there's kind of a physics smell to the breakthroughs that led to these systems called Boltzmann machines and Hopfield networks that Geoff Hinton and Hopfield got the prize for. They're very physics-y, and they look like the same exact objects that physicists study.
Matt Turck [41:33] That's super. Just to play it back, part of what you were saying is the Nobel Prize going to AI is less a function of AI sort of eating everything, but almost like a political thing at the Nobel Academy, or whatever it is, that they should have had a prize for computer science and they don't have it. And now it looks kind of silly because AI is the fastest-moving field in the world. Therefore, they're sort of retrofitting it.
Misha Laskin [41:59] Exactly. And I don't mean it in a negative way. But yeah, I don't think that, as a physicist, when I look at, again, these objects, Boltzmann machines, Hopfield networks, very fundamental objects in the development of AI, even though they're not really even used today, some concepts from them are, those things haven't permeated physics in any way whatsoever. But they just look like objects that a physicist would study.
Matt Turck [42:05] Mm-hmm.
Growing Up in a Science-Driven Desert Town
Misha Laskin [42:07] Kind of mathematically look very similar.
Matt Turck [42:08] Physics-y.
Misha Laskin [42:09] Yeah. Okay.
Matt Turck [42:16] All right, so you evolved from physics to AI, and then what was the next step?
Misha Laskin [42:35] Had an interim where I started a small startup that went through Y Combinator, and it was basically doing machine learning prediction for inventory management. Really felt like I had to get on the frontier of AI research. I felt that was going to be where a lot of scientific impact accumulates. And so I ended up joining UC Berkeley as a postdoc, where I worked in this lab called the Pieter Abbeel Lab, which is one of these great labs for reinforcement learning and what was called unsupervised learning research, which is basically—large language models and diffusion models are the biggest outputs of unsupervised learning.
Misha Laskin [43:27] I didn't realize that at the time, but that was, in some sense, a—I wouldn't say miracle year, but it was a special year in that lab, given the people who were in it. A large portion of that lab ended up going to start impactful companies or being scientists doing very impactful work in large labs. To give you a sense, the first people I worked with were Arvind Srinivas, who now runs Perplexity, and Dennis Yarats, his co-founder. We worked together on some papers there.
Misha Laskin [43:56] Jonathan Ho, who is one of the inventors of diffusion models, was there and invented diffusion—the big paper that made them break through there. And this guy named Ajay Jain, who was on that paper, and they started Ideogram and Genmo, which are two startups in the video gen and image gen space. Deepak Pathak, who is the founder of a company called Skild AI, which is one of the premier robotics companies, was there.
Matt Turck [43:58] What year was this? What rough time period?
Misha Laskin [43:59] 2020.
Matt Turck [44:00] 2020?
Misha Laskin [44:11] Yeah, it was—now when I look back at it, it was a pretty incredible group of people, like Aditya Grover, who started Inception, which is a company that does diffusion models for coding. I was just thinking, I saw this guy's tweet, his name is Kevin Lu, the other day, who was in the lab, and he was an undergrad then, and most recently was leading a lot of the work for the mini models at OpenAI.
Misha Laskin [44:43] So just an incredible group of people at that time. Not obvious at all then. I mean, the research people were doing was very interesting, but I would never have predicted that so many companies would have come out of there.
Matt Turck [44:45] And the next step after that was DeepMind.
Misha Laskin [45:05] And then I went to DeepMind. At the time, I was really interested in Toronto and New York. So I joined the group of this researcher named Volodymyr Mnih, who was largely credited with starting the field of deep RL. He was the first author of the Deep Q-Networks paper, which was the paper that got neural networks to play Atari. And his first set of papers actually largely defined deep reinforcement learning as a field then, and were basically DeepMind's claim to fame for a very long time.
Misha Laskin [45:32] And so I joined his group. The team we built together was called the General Agents team. And so the whole point was to do research for us to figure out: how do we build general agents? I think it was much more opaque then than now. And the big problem we were trying to solve is what people called, and still do, unsupervised reinforcement learning, which is really: how do you train reinforcement learning systems that are capable of assigning their own rewards?
Misha Laskin [46:13] If you don't have rewards without supervision, in the same way that you can give things some rewards, but kids and animals, when you look at them, they learn a lot in an unsupervised way. They interact with their environments without—no one is telling them to. And so we were thinking about how do we teach reinforcement learning systems in this way? I actually think the subject is coming back in vogue in the age of language models, with the question of how do you do pre-training-scale reinforcement learning?
Misha Laskin [46:43] How do you generate a lot of synthetic data if you don't have explicit rewards? I think it's actually a really interesting question now again, but that was the general agenda of what I joined to study. And the short of what ended up happening was that language models started working. And once they started working, that really changed, I think, my entire perspective on what problems matter and what didn't. Because a lot of the problems that we thought were fundamental problems were solved in this brute-force way for us.
Misha Laskin [47:13] Gemini 1.5, and then obviously 2 and so forth. And I joined with my co-founder. My co-founder, Ioannis Antonoglou, was leading the reinforcement learning team, the RLHF team. I joined his team and led a lot of the work for training reward models for Gemini and implementing the algorithms and so forth. And it was a very exciting time when I'd say a group of 10 to 20 people were—that was basically—all the people doing the RLHF work there.
Matt Turck [47:35] And when was the decision to leave and start a company? And what was the thinking?
Misha Laskin [48:03] After Gemini 1.5, we realized that language models crossed this threshold of utility where they're no longer research objects. They're going to be very useful. This was early 2024. We realized that the ingredients were in place to build a superintelligence. We felt that everything was there. There was one more piece to solve of going from RLHF to making reinforcement learning work. And that basically happened over the last year with reasoning models. So we felt that that would happen. Then the question was: superintelligence for what?
Misha Laskin [48:31] That was basically—we felt that you can't answer this question in the abstract by being a researcher that's really far away from product and customers. You really had to go in and define what that means from a product vision and what problem you're trying to solve perspective. It's like, we're not interested in building a superintelligence that will be superintelligence in mathematical Olympiads. And the difference between this era of reinforcement learning and the previous era of pre-training is that when you did pre-training, you made the models generally better at everything.
Misha Laskin [49:02] Reinforcement learning is much more jagged, right? It makes them good at what you want them to be good at. So just because you made them good at competitive code, that improves the general coding capabilities. But that does not mean that you'll have not even a superintelligence, but even just a useful intelligence for software engineering code. And I think an example of that is, I think Anthropic has done a really good job of building models that are meant for users of their products rather than benchmarks.
Misha Laskin [49:39] When I look at academic benchmarks, the Claude models are consistently worse than whatever else is out there, oftentimes not even close. They're consistently worse. And yet, from a user perspective, they're consistently better. Something has to explain that. And I think the explanation is that when you train large language models with reinforcement learning, they become jagged in the sense that they become good at what you wanted them to be good at.
Matt Turck [50:11] And there are some generalization capabilities, but they're much weaker than people think, which is a little counter to the narrative that you hear a lot, which is that generalization is always going to win. And I guess, I don't know if that's true to it or a bastardization thereof, but like Rich Sutton's Bitter Lesson. And so what you're saying is sort of not the opposite, but that the solution is a combination of generalization and specialization. Is that fair?
Misha Laskin [50:37] Well, I think the Bitter Lesson actually doesn't say anything about generalization. The Bitter Lesson says that the systems that we should be thinking of and building are ones that scale well with search and compute. That's kind of it. And so what he's saying is, if models are limited today—this was actually, basically, the lesson was for researchers, but I think it's for product builders as well—if you're building your product with the assumption that these models are going to stay at their current intelligence level, and you make a bunch of hacks around your product to overcome those things, then in the next iteration of models, a lot of the hacks that you put in place will probably be Bitter-Lessoned.
Misha Laskin [51:20] And that was scientific researchers were seeing these kinds of limitations of models and plugging them in with kind of temporary hacks that would make them better at certain benchmarks. So the same lesson translates there. But the lesson is more to build systems that are good at soaking up compute and scale well with search.
Matt Turck [51:27] Yeah, we'll put that in the show notes. I think that's probably one of the most often misquoted blog posts in the history of AI.
Misha Laskin [51:47] Well, I think what's happening when we think of generalization, what's really happening is that if your training distribution is everything, then your test distribution just falls in your training distribution, and you have generalization. Maybe one point of view that I don't think that many people share, but I do think there will be a general superintelligence. But I think that it won't be one lab that has built it, but it'll be kind of the plurality, like the collection of all intelligences, will be a general superintelligence.
Misha Laskin [52:23] Because if you think of pre-training as kind of the soil or the substrate from which you can now grow superintelligence in various categories—medical superintelligence, organizational superintelligence, superintelligence for scientists and math—these are all different types of superintelligence. Some of them you can merge into one model. But I think that there'll be, if you think of different superintelligent plants growing from the substrate, yes, like a frontier lab will be able to capture some of them. But I think that there'll be new frontier labs built that build out, sort of grow other plants, and that the collection of this garden is going to be a general superintelligence rather than one company going in and growing all the plants.
Misha Laskin [52:53] I think from a research and compute perspective, that's possible. But from a product perspective, if you want to build organizational superintelligence, you have to go and integrate with all these customers, and you have to have solution engineers that support them, and you have to have salespeople that support them. And in this kind of rosy picture of a researcher who just trains models and hopes the model is a superintelligence, I just don't think that the world will play out that way, because it's meaningless if it's not coupled to a product and its deployment.
Building General Agents at DeepMind
Matt Turck [53:31] Double-clicking on the product today, what is the reality of something like an autonomous coding agent? How good is the state of the art right now versus what hopefully it will be in the future? Are we in the teens in terms of SWE-Bench? Are we higher than that for certain tasks? Where does it all land currently?
Misha Laskin [53:50] On the benchmarks, a quick side comment: I think that SWE-Bench is getting to saturation. And it's funny that even though you have these numbers, like 70% or something like that on SWE-Bench, those coding models are good, but they're not solving 70% of engineering tasks. So there's a sort of benchmark-to-real-world-problem misalignment, which is always going to be the case when your benchmark is not the actual thing that customers are using it for.
Misha Laskin [54:20] If there's one thing, I had fairly aggressive timelines on progress in my mind, and I would say things have moved faster even than I would have expected. I think that we've gone from autocomplete engines to things that are kind of semi-autonomous to now things that, for junior tasks, can just do them autonomously. It's pretty incredible. There is, in some sense, we are probably at an L4 kind of junior engineer level of autonomy, which is pretty incredible.
Matt Turck [54:32] And autonomy means 100% success, no need for code review?
Misha Laskin [54:55] You still need to do code review in the way that you would do with an L4 engineer, but at a level of reliability where there will be some nits that you pick off, like you do with a normal engineer, but the thing they gave you is useful enough for you to review it in the first place, as opposed to, "This is just garbage, and why am I spending time reviewing this code?" So I think for junior code, little UI changes and small things kind of here and there, of which there are a lot.
Misha Laskin [55:26] So it's very useful. We're probably at L4 as a field. I think that over the next couple of years, the code generation ability will continue improving, and we'll have things that are quite intelligent. I don't know where exactly I'd place them because I think they'll be really good at generating code when you give them all the specs, all the requirements, exactly what needs to be built. But the whole job of a staff-level and above engineer is figuring out the requirements in the first place.
Misha Laskin [56:02] So that's kind of 70% of their time is spent on that: designing stuff, figuring out what the requirements are, planning in advance. Like, if I build this, will it conflict with these things? Soliciting information from other teammates. That's really what a staff engineer and above does. And then the implementation part is usually the sort of straightforward part. Okay, you spend like 20% of your time actually writing code and implementing things. And so the way things are progressing, I think we'll have very capable agents at the implementation level once you give them something to do that's really concrete and specific.
Misha Laskin [56:21] But if you don't solve this kind of contextual gathering or build out the contextual core, I don't think they'll be at that staff level as a whole.
Matt Turck [56:35] But the L9 with memory, not the L9 with amnesia that we were talking about earlier, feels like a tractable, near-term kind of problem to solve from your perspective. Like, we're well on our way there.
Founding Reflection AI After DeepMind
Misha Laskin [57:01] Yes. So I think it's a very hard, very tractable problem. And the combination of this LLM with amnesia and the LLM's context core, together, will become the principal-level engineer, an AI engineer. And so I actually think that that's not too far away. That's, I would say, a couple of years away.
Matt Turck [57:31] So when you say, rewriting back to the beginning of the conversation, that you're starting with a coding agent, then is the idea that a lot of those principles that you described can be horizontalized across the company? So the institutional memory, which I find a fascinating concept that you can plug in, that could be the coding institutional memory, but next it becomes the marketing, the product, the HR institutional memory.
Misha Laskin [57:59] Yeah, that's right. It's kind of, at that point, from a build-out perspective, very similar, right? You're now on-prem or in the VPC of an enterprise customer. You have a centralized kind of knowledge around their code, and you've already centralized knowledge around other tools for them as well that are immediately adjacent to the next thing, right? If you're centralizing knowledge from Jira, that's both engineering- and product-management-adjacent, right? So I think at that point it just becomes adding other tools, other kinds of integrations, based on where you're seeing pull from the enterprise, and then enabling the ability to act on the user's behalf when they want to.
Misha Laskin [58:39] So instead of just being able to ask questions, enabling the actual agent to go and do stuff for them. So I think that that's going to come sooner than later, right? In the coding space, I think once you have that contextual core, you can integrate it into existing coding products, right? You sort of, like, if we think about the other coding products as being these autonomous software engineers with amnesia, you can fill that gap for them. Obviously, you can build out one of your own as well, but I think it ends up being kind of a notion of customer choice, right?
Product-Driven Superintelligence: Why It Matters
Misha Laskin [58:58] You want the overall solution to be the best for the customer. And you're, as a company, focused on what you believe is a fundamental building block of enabling superintelligence.
Matt Turck [59:33] So maybe zooming out to close, I'm curious on a few thoughts about the reality of building an AI startup today, I guess, from a talent perspective to start with. As we were saying a few minutes earlier, we're in this weird moment where ridiculous amounts of money are being offered to talent to move from one company to the other. How does one recruit and keep talent in this environment when you're very impressive, obviously, but still a small startup?
Misha Laskin [1:00:01] I think the entire makeup of the research team was earning a lot of money at big labs, basically, obviously, Ioannis and myself included. The thing to remember is that a lot of people get into this field because they're scientists at heart, or they're builders at heart. And so there is the financial element. You certainly need to pay enough, be generous enough, where that's not really top of mind. But people really care about discovering the next frontier and the next breakthrough, right?
Misha Laskin [1:00:30] The most exciting time to be in an AI lab is before it's obviously the frontier lab. I think the most exciting time at DeepMind was building Deep Q-networks and building AlphaGo. Those were the times, I think, where DeepMind was in its golden days. And similarly with OpenAI, it was building out the GPT series of models, like at the GPT-1, 2, 3 stage. And I think for Anthropic, it was really in the Claude 1 and 2 stage, where obviously now they're reaping the benefits of that breakthrough work that was done there.
Misha Laskin [1:01:02] We tend to attract people who have that kind of internal drive in them of wanting to be part of that next story, because they're already at a big lab, or they could join a big lab, and they'll always be able to. I mean, we'll see when ASI comes around, but that's not really that scarce of an opportunity.
Matt Turck [1:01:02] Yeah.
Misha Laskin [1:01:26] When you actually look at which startups are out there that have the clarity and the team and potential of starting a new frontier lab, there are not that many. So there are actually a lot more spots at the big labs than there are at startups that have a shot at this. So I think that people end up, in a sense, self-selecting. We win over candidates over OpenAI, Anthropic, Meta, DeepMind regularly. Obviously, they get a lot more equity in this company as a percentage of its ownership.
Misha Laskin [1:01:42] And if they do back-of-the-napkin math of if they had joined Anthropic at this stage and got that percentage ownership, what it would be worth now, it would be absolutely generational.
Matt Turck [1:01:43] Yeah.
Misha Laskin [1:02:12] So that tends to be, I think, if you don't have a good kernel of the initial team, if it's not strong enough, then it becomes very hard. But if you have a very strong initial team, and people see that potential for breakthroughs, then you become, in a sense, a scarce option because there are not that many places where you can do this. And later this year, we'll be shipping things that I don't think anyone ever thought a startup could do. I think that we're going to be shipping some things on the research side that I think everyone thinks you need to be a giant lab with 100,000 GPUs to do.
The State of Autonomous Coding Agents
Misha Laskin [1:02:26] And I think it'll be quite interesting and surprising.
Matt Turck [1:02:46] On the product front, again, to the discussion about product versus research, you guys are super deep, PhD-type, world-class AI researchers, but in a context precisely where you want to build product, was it part of the core team to bring in people that would bring product, or how did you think about it?
Misha Laskin [1:03:12] Yeah, we've built out, I guess, the company is—we kind of think about half product, half research. And so we've built out a research team, we've built out a product team. And then there's, I'd say, the majority of the makeup of the company is probably two-thirds of it is people who have research backgrounds at some of the big labs. Of those people, a bunch of them are kind of in this role that is between research and product. For example, the design of the agent, design research—that's very between research and product—or evaluations.
Misha Laskin [1:03:38] Like, what are you evaluating your models to be good at? That typically is something that's just on research. For us, it's kind of a cross-functional, end-to-end thing. The data that you're generating, the synthetic data that you're generating to train your models, that also cuts across all those things. So in a sense, I think it attracts, well, maybe people similar to Ioannis and myself that came into this and just wanted to be—we just did not want to maximize another academic benchmark.
Misha Laskin [1:04:14] We just wanted to solve real problems and have real evaluations. And so for those people, this ends up being a really good place. I think for people who would much rather kind of sit in a known entity and really focus on some specific piece of work of training the model, because they're so big that when you enter them, you're kind of given, okay, this is the sliver that you own. But that's really interesting nonetheless, because you kind of become a craftsperson.
What's Next for Reflection AI
Misha Laskin [1:04:30] So I think that's actually a very important skill to pick up. But if that's where people are in their lives, then, yeah, I think the big lab is definitely a better option for them.
Matt Turck [1:05:01] And you've raised a bunch already, I think like $125 million, $130 million, whatever the number in the press is. Does capital matter as much for a company like yours, thinking about moats in AI today? There certainly is a well-publicized capital race for certain types of companies. Does your generation of startup and your type of startup require as much?
Misha Laskin [1:05:24] Capital matters a lot. I think the difference is that you can't operate at 100x less capital than a frontier lab, but you can operate at, say, 10x—like an order of magnitude less capital—when you're really focused. So I think that capital matters a lot, and it really is determined by when you're ready to scale up your GPU count. That's basically it, right? I mean, there's obviously headcount and data and so forth, but the primary cost for any of these companies is their GPU expenditures.
Misha Laskin [1:05:46] So that's sort of, you kind of raise capital commensurate to when you're ready to scale to the next stage. And so I think capital matters a lot, but you can be a lot more efficient than traditionally frontier labs have been.
Matt Turck [1:06:07] All right, well, we covered a bunch, from the path to superintelligence to Asimov to your background to what you're building and what you're building next. I'm excited for this announcement, Asimov, and then what you alluded to that's coming on the research side in the next few months. Thank you so much for being here. This was terrific. Thank you.
Misha Laskin [1:06:08] Yeah, thank you, Matt. This was fun.
Matt Turck [1:06:29] Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.