Mistral AI vs. Silicon Valley: The Rise of Sovereign AI
The MAD Podcast with Matt Turck · with Timothée Lacroix, Co-Founder & CTO, Mistral AI
Timothée Lacroix is the Co-Founder & CTO at Mistral AI. We cover why Mistral builds data centers for stable training on thousands of GPUs, why enterprise AI succeeds through governed workflows rather than autonomous agents, and why the real ROI arrives after companies connect and safely reuse their internal context.
Chapters
- 1:27 — Mistral vs. The World: From Research Lab to Sovereign Power
- 3:48 — Inside Mistral Compute: Building an 18,000 GPU Cluster
- 8:42 — The Trillion-Dollar Question: Competing Without a Big Tech Parent
- 10:37 — The Reality of Enterprise AI: Escaping "POC Purgatory"
- 15:06 — Why Mistral Hires Forward Deployed Engineers (FDEs)
- 16:57 — The Contrarian Take: Why "Agents" are just "Workflows"
- 19:35 — Trust > Autonomy: The Truth About Agent Reliability
- 21:26 — The Missing Stack: Governance and Versioning for AI
- 26:24 — When Will AI Actually Work? (The 2026 Timeline)
- 30:33 — Beyond Chat: The "Banger" Sovereign Use Cases
- 35:46 — Mistral 3 Architecture: Mixture of Experts vs. Dense
- 43:12 — Synthetic Data & The Post-Training Bottleneck
- 45:12 — Reasoning Models: Why "Thinking" is Just Tool Use
- 46:22 — Launching DevStral 2 and the Vibe CLI
- 50:49 — Engineering Lessons: How to Build Frontier AI Efficiently
- 56:08 — Timothée’s View on AGI & The Future of Intelligence
Transcript
Mistral vs. The World: From Research Lab to Sovereign Power
Matt Turck [1:26] Hey Timothée, welcome.
Timothée Lacroix [1:27] Hi.
Matt Turck [2:01] So as I was prepping for this, I was struck by how much has been going on at Mistral over the last few months. I think most people probably know Mistral as a provider of open-source models. It seems that you guys evolved from an AI lab to more of a full-stack solution focused on enterprise and sovereign customers. Seven billion Series C led by ASML, seven billion post-money valuation, you launched a bunch of models, which we're going to talk about. Is the big vision behind all of this that enterprises and sovereign states are going to need their own AI infrastructure, and Mistral is going to be the provider?
Timothée Lacroix [2:51] So the big vision has been evolving. And as you stated, we started as a company that built models because, with Arthur and Guillaume, this was what we knew how to do at the start. The premise on which we built Mistral AI was immediately solving for enterprise needs. And we started with open-weights models. After this, and working with enterprises, we realized the need for basically the rest of the stack. So we built the serving platform because infrastructure was needed. And then all of the tooling around it was also something that we saw was missing.
Timothée Lacroix [3:24] More than the tooling, it also requires a lot of work and expertise still to get deep into an enterprise workflow and really help that transformation. And so we built that FDE function. And more recently, with Mistral Compute, we're going a bit lower in the stack as well. So we've done all of this because it was required for enterprise success while still continuing on our models journey. All of this stack being modular is really important to us, as it gives full control to enterprises and our clients as to which part of the stack they decide to own and control, which is maybe more involved, or that they decide to have serverless, or basically this modularity that we like.
Inside Mistral Compute: Building an 18,000 GPU Cluster
Matt Turck [4:11] All right. So let's take some of those modular components in order. Let's start with Mistral Compute. So that was a big announcement, I guess, in June of 2025, putting a partnership with NVIDIA to help with this effort. What's the current status? Is that live yet? Are you building it? How does one go about building data centers or leveraging data centers in Europe?
Timothée Lacroix [4:32] Maybe first to go into the reasons why we decided to start building our own data centers. We tried a lot of different partners over the years, and we realized that our use of AI compute for large-scale training was not necessarily well understood by a lot of providers, and our need for stability especially. Like, when you run inference on a few GPUs or when you run small-scale trainings on hundreds of GPUs, margin for error is a lot larger than when you run trainings on thousands of GPUs at the same time.
Timothée Lacroix [5:16] And so, to address this need for stability, we saw a way for us to basically build our own data centers and maintain them with our understanding of what quality looks like. And so that was why we launched Mistral Compute. And when we decided to do it, we also realized, well, maybe others will benefit from it. We launched into a bigger development than what was previously intended. And so this was announced in June, as you said. Since then, the building of the facility has progressed quite well.
Timothée Lacroix [5:48] It's in the south of Paris, and we are right now running through the stabilization of the first tranche. So it's quite a large data center, so delivery doesn't happen in one day. And the first part of this data center is something that we are working on as we speak. We have a few jobs running, and we're fine-tuning basically all of the last things to run at speed and with the right stability.
Matt Turck [6:00] Okay, great. And did I understand correctly, it's going to be for your customers and your own needs around training? But also you'll be providing it as a service to others in Europe and beyond?
Timothée Lacroix [6:11] Yeah, exactly. So we will use part of that capacity for ourselves as one of our training clusters, but we will also provide a managed Kubernetes and managed Slurm stack on top.
Matt Turck [6:29] Okay. Any lessons learned so far? I mean, as you said, you guys come from a very deep background in AI and AI research. It's a whole different thing to build a whole data center facility. How have you gone about it, and what are some things that surprised you and any lessons so far?
Timothée Lacroix [6:54] As with most new experiences as a founder, I relied on the knowledge of others. And so I was lucky to have a few seasoned HPC experts and a lot of cloud software experts as well to build that solution. For me personally, and it's one of the things I love about my position at Mistral, is that I get to discover so many new things and so many new problems I hadn't thought possible. Having to learn all of the different parts of building a data center, all of the different trades that you have to coordinate, all of the potential synchronization between all of the different trades.
Timothée Lacroix [7:32] I mean, it's a huge building. It involves hundreds of people working on it. Then, when you stand up the thing, you have to question what works. You have to filter through the blades that are faulty. It's just an entire new area of work where I get to see experts in their field go through things and try to explain to me what their daily work is. It's always fascinating to see an expert in his field do something that you don't know how to do.
Timothée Lacroix [8:03] I think the logistics of it and the timelines are also quite different from what I'm usually dealing with in software and research. For new capacity to be built, you have to plan around having energy available, you have to plan for the space to be available and on time. And so it's a lot more long-term planning than a few software features.
Matt Turck [8:07] How do you guys go about power, since you mentioned energy?
Timothée Lacroix [8:39] In what we've been doing in Europe so far, it hasn't been a huge blocker, although there are constraints. I think the grid in various parts of Europe is not necessarily easily extensible. I know it's an issue in France. A lot of the sites are contended, so we'll see how it all develops. We are lucky in Europe to have very clean and affordable energy, either with green energy in the Nordics and nuclear in France. So it's been relatively okay for us today.
The Trillion-Dollar Question: Competing Without a Big Tech Parent
Matt Turck [9:12] As you describe this, what comes to mind is the gigantic amounts of money that are being invested in the US around data centers. How do you guys go about that from a financing standpoint? And perhaps even more, taking a step back, if you think about the race between the big AI labs globally, whether that's the OpenAIs and Anthropics of the world and xAI, it seems that all of them are affiliated with a gigantic pocket of money somewhere. Obviously, there's Gemini and Google to add to the list, and Meta.
Matt Turck [9:32] I'm just curious, where do you guys stand on that? You have a bunch of partnerships with SAP and NVIDIA, but you don't have one of these gigantic companies on your cap table. So how do you think about competing in that general context?
Timothée Lacroix [10:08] So with those companies, the hyperscalers, there are two parts to the game, and we've played the partnership part quite well with them, and we're integrated within Google Vertex AI, Amazon Bedrock, and Azure AI Studio. And that is the choice that we've made in terms of having access to gigantic pockets of money. We've been focused on efficiency from the start, and I think we've done quite well at building models that are competitive with the investments that we've put in. For us, it's important to build the company as efficiently as we can, and I deeply believe that with the capabilities that we have today in the models, there is so much to be unlocked in enterprise.
The Reality of Enterprise AI: Escaping "POC Purgatory"
Timothée Lacroix [10:38] That I don't think my main focus today would be going into the gigawatts of power. We still need to build so much with our clients and unlock so much value with the capacities that we have.
Matt Turck [10:55] All right, so let's go into the enterprise reality of all of this. So if I'm an enterprise or if I'm a sovereign and I want to deploy a Mistral open-source model, what is it that I do these days with everything that you've built?
Timothée Lacroix [11:22] The way we work with enterprises, as you mentioned, we have a few of our models that are open source and Apache, and all of our clients are welcome to use them as they need. What we have seen in terms of success is that, given the current stack, it still requires a lot of expertise to manage to get to actual value and things that go to production, basically. The way we interact is that we usually stand up our Mistral AI Studio, which is our platform, and we can deploy all of our stack on the client's choice of deployment methods.
Timothée Lacroix [12:00] So it can be on-prem, it can be on their VPC, it can be in several places. The reason we do this is that it lets clients build where their data is without having to shuffle things around, which, as I've learned as a CTO, is something that you don't want to do ever because it raises a lot of questions and it's quite a stressful thing to do. So once this is deployed, we then work with the business units to understand where their pain points are.
Timothée Lacroix [12:43] Sometimes it's knowledge management, and I think it's the most well-known use case from the outside of the enterprise world. But it's also around automating core workflows for the enterprise. It's some tooling that you wouldn't expect, where one thing that we've done is around code modernization, where you turn a bunch of Excel sheets into an actual Python app. And if you have many, many of those sheets, then potentially you want to use AI for this. So once the infrastructure is built, then we basically look for what's the most valuable to the customer, and we start accruing value inside a stack of AI assets that then accelerates all of the other developments with that customer.
Matt Turck [13:09] And is part of the idea that you do actual model work at the customer and for the customers, in particular fine-tuning?
Timothée Lacroix [13:31] Yes, we customize in various ways. So we have done continued pre-training, and this is most useful when you want to change the capabilities of a model more deeply. So we've done this to sometimes change the mix of languages in a model to get something that's a lot better at Southeast Asian languages, for example. Or you could require this if your internal data, which doesn't happen on the public web, is something that's so new that you need a large amount of tokens to get a model that understands it and becomes fluent with it.
Timothée Lacroix [14:11] So we do these kinds of continued pre-training. Fine-tuning, we also like, and this is more for an efficiency reason. When you get to smaller models, you have to make trade-offs. The models won't be as good in their knowledge of the world, and so you lose a lot of things. You have to focus on what you really care about. And so this is typically important if you want really fast, really cheap models that will be really good at a specific task.
Timothée Lacroix [14:40] It's also useful if you want models that run on the edge that get very, very tiny. And so for all of these, fine-tuning is a tool of choice. Another reason to do fine-tuning can be to adapt to data that's not necessarily massive, but that's also not available on the web. So typically in coding, what happens is that you will have massive codebases, sometimes accrued over decades, that the model will need to be able to work with in terms of having AI deployed on it, typically.
Timothée Lacroix [15:04] And so being able to come in, not move the codebase, and train an actual coding agent for that codebase is really powerful as well.
Why Mistral Hires Forward Deployed Engineers (FDEs)
Matt Turck [15:09] And who does all of this? You have evolved towards an FDE model?
Timothée Lacroix [15:41] So we have indeed a large FDE section. It's a mix of software and FDEs, and we split our FDEs into what we call AI engineers and applied scientists. And so applied scientists will tend to use the tools that we've just talked about, so fine-tuning, continued pre-training, and the like, whereas AI engineers will focus more on adaptation to the enterprise environment and figuring out what workflows to automate and all of this. They work with the customers to make sure that the use cases are indeed providing value and going to production, but it's also a fantastic way for us to understand what matters in an enterprise context and be faster at building the right platform.
Matt Turck [16:16] And again, those customers are the kind of customers for whom customization and privacy are essential. How do you position, again, the OpenAIs and Anthropics of the world that are going very hard at the enterprise? Is that data sovereignty? Is that customization?
Timothée Lacroix [16:43] The term we use is control. The value that we see is both in our expertise and the software stack that we provide. The software stack, once deployed, is in the hands of our customers, and they can change it. They can add to it. They own model changes that we make. And I think it's really important as a customer to consider that your expertise and what makes your company valuable stays yours. And so, in working with us and building—because it takes effort to build an AI advantage today—and so having this effort built into something that you own is, I think, a choice that makes sense.
The Contrarian Take: Why "Agents" are just "Workflows"
Matt Turck [17:10] Let's talk about agents, obviously part of the overall effort at Mistral. How does that work? How do you build an agent? And what key use cases have you seen so far?
Timothée Lacroix [17:37] Personally, I think I've moved from agents to workflows, which is, I guess, an abstraction on top. So agents are, I think, the building blocks where you have a given expected input, a set of tools, and you have a goal that you want to reach. The set of inputs that we've enabled are images, text, and audio. When you build an agent, to me, it's really important that you build it on a focused task with a dataset that you understand and that you can iterate on and improve.
Timothée Lacroix [18:26] What we see in enterprise is rarely things that are solved with agents because that's not necessarily where you would expect an FDE to be most useful. Those ideally would be built on our platform by the customers directly. Where there is more value is in more complex workflows, where you will have several agents interact through a workflow to automate something slightly more complex. And so that's what we've been focusing on.
Matt Turck [18:27] What would be an example?
Timothée Lacroix [18:59] An example is something that we've built with the shipping company CMA CGM, where we've automated the container release process. And so it's a use case where, I don't know how familiar you are with shipping. I wasn't at first, but a container reaches a port and you have to arbitrate, probably in English, some decision. A decision has to be made that this container is ready for release to the next person on the line to handle this container. And so there are lots of checks that need to be run and data to be accessed in the backend before that decision is made.
Timothée Lacroix [19:33] So as you can imagine, some of those containers are extremely valuable, and you can't really afford a mistake. And so what we've done in this case is an application that's integrated into how these harbor workers work, and it automates a lot of the manual work that they did to check the data, and they make the final decision given all of the evidence.
Trust > Autonomy: The Truth About Agent Reliability
Matt Turck [19:47] Okay, this is super interesting. Obviously, the key question about agents these days, especially when they are combined into workflows, is the question of autonomy. How do you guys think about it? How autonomous are those agents?
Timothée Lacroix [20:15] I don't know if it's the way I think about it. To me, the better question usually is how much you trust the agents. And there are a few dimensions around this. What worries me when building those kinds of workflows is that typically, if you want the value to accrue, and if you want to build faster and faster, the more workflows that you build, what you will want to do is reuse assets and make them reusable by others. As soon as you do this with agents, you then start to ask the question: well, this agent has access to some data that is privileged, but maybe this other agent is publishing it to something that's public.
Timothée Lacroix [20:57] You might have governance concerns where some agent is acting on something very critical, and you don't necessarily know that the data that it got has been approved, or something like this. It's really a new way to develop, where the parts of your workflows have to be trusted. Each of them, to be trusted, requires quite a lot of tooling and quite a lot of observability to get confidence and to basically enable this at scale in an enterprise. So the question that you're asking about autonomy, to me, this is something that I see happening when I vibe code.
The Missing Stack: Governance and Versioning for AI
Timothée Lacroix [21:26] Sure, longer-running tasks and making and improving on this is going to be critical, and we're working on it daily. But today, the problems that we're solving on the software side of things are really about how you trust what you've built, how you improve it, and how you allow an entire company to build on it with confidence.
Matt Turck [21:39] Maybe describe some of the things that you guys have built in AI Studio around governance, as you mentioned, and traceability and registry, all the things. What are the key components of a modern agent suite?
Timothée Lacroix [22:11] So workflows, as I mentioned, is something that we've worked a lot on with our customers, and it's not GA yet. So look out for this sometime in the future. But it's also one of the benefits of working with enterprises. We can have a lot of design partners, and once we're confident with the solution, we make it GA. So a workflow solution is critical. Workflows are built on various model capabilities, so vision, audio, text, and reasoning. It is important to have a registry of connectors and MCPs.
Timothée Lacroix [22:37] And so for this, we have our connectors. Observability is an area where we're still working on. It's important for me to be able to iterate and really define precisely what an agent does, control each of its goals, and see how it's progressing, being able to maintain evaluations and build on them. What is difficult in this entire sea of complexity is that you also have to maintain proper versioning and tagging and think about how you're going to deploy and improve upon what you've built.
Timothée Lacroix [23:13] So let's say you've built a KYC workflow based on a lot of agents and models that Mistral has released in the past. Then a few months pass and there are new sets of models that are out. Maybe you can simplify that workflow. Maybe the next Mistral 4 is good enough that you can factor out a few agents. Basically, what you need to be able to do is create a new agent, run it on the same set of inputs and outputs, and control that you haven't broken anything, and then deploy it in the wild.
Timothée Lacroix [23:39] All of this software suite, basically, which has been built for software development over years, I feel isn't there yet in the AI world, and that's what we're building.
Matt Turck [24:02] As I'm sure you've seen, for the last few weeks in startup and venture circles, there's been this whole idea of the context graph as an infrastructure that made the rounds. Is that something that you think about, a layer that would basically enable one to know how the agents made a decision and how those decisions relate to one another?
Timothée Lacroix [24:28] I've seen this indeed, and I think there are two levels to that discussion. The part that you mentioned at the end, where it's interesting to know how an agent came to a decision or an action, the game is really to understand how a human agent really made this decision.
Matt Turck [24:28] Yes.
Timothée Lacroix [24:56] It's understanding how an enterprise does what it does, and it's certainly interesting. What keeps me up at night and what I really want to solve first is just the basic idea of gathering a workable enterprise context. Right now, with any model and with a lot of effort, you will be able to get some connections to tools, and you will ask questions and your agent will do a bunch of things. It will realize that by doing five API calls and three joins, I can probably get what Timothée asked. Immediately, what should happen is that all of that discovery and all of that intelligence should be stored somewhere to be reused.
Timothée Lacroix [25:46] It's not really how things happen. It's just basic knowledge about what the infrastructure of the company is. So knowing where the tables are, what they contain, how they're joined. So all of this is compute that should be amortized, basically. And to me, it's really the entire game with the context engine, as we call it internally, is to be in a setup where, over time, knowledge of the company and the context that's available to the agent accrues and is maintained. The second-order thing of, how was that decision reached?
Timothée Lacroix [26:19] Sure, it's going to be super interesting and it's important. But right now, I feel we're not even in a place where it's easy for an enterprise to have any worker in it be able to build an agent that has access to the right context. For this to happen, you have huge data privacy concerns. If you want this to be efficient, you need to give access to the agent system to the entire data of your enterprise. And there are going to be RBACs everywhere, and you need to make this safe.
When Will AI Actually Work? (The 2026 Timeline)
Matt Turck [26:36] Speaking of which, what's the current reality of enterprise deployments of generative AI from your perspective? Just listening to some of the concerns, it seems like we're very early.
Timothée Lacroix [26:56] To me, we are still in the building phase. And I think it's kind of the frustrating thing for enterprises, is that when you come to a chat assistant, you feel that it's magic and it's all going to work. But as with most things that have value in life, there is still work to be done to get to them. And so most of the enterprise value of AI will happen once you've gone through that first building phase of just setting up all of the machinery.
Timothée Lacroix [27:26] You've got to set up all of the connections. You've got to make all of that data available. And the reality is, even despite a lot of work recently to make data more available in enterprise, it's still not easily available in the format and at the scale that we need for the true ROI of AI to happen. And so, when we come in, there is still that phase of work that is just work to connect everything and then be able to build on it.
Matt Turck [27:41] So do you think we are years away from generative AI actually being deployed in the enterprise?
Timothée Lacroix [27:53] Not years. I think a year, singular. It's also, to be fair to us, we've started working—I mean, the company started two years ago. And so most of our—
Matt Turck [28:01] It's a good reminder, right? It's a good reminder that you have done all of this and the company was started in, yeah, June '23, right, if I recall.
Timothée Lacroix [28:25] Yeah. And so for most of our clients, we started working with them recently. The tooling for everyone is still in its infancy. And so I hope that the tooling will stabilize and I hope that we will have true value. True value to me is really, okay, we've gone through that first phase of building connections, and now employees of that enterprise are able to use everything that we've built. Right now, I think we're in a phase where we build siloed things because we're scared of data going through walls and everything.
Timothée Lacroix [28:43] And so, to me, the real success is when you're confident enough to give all of that control back to the company's employees at large and they start really building on it.
Matt Turck [29:06] You're talking about Mistral in particular, but the industry in general, right? Do I understand this correctly? Because obviously that's the big question, right? Are we all collectively building this whole thing—and data centers and models—and pouring billions? And I think it's pretty clear that from a personal use case or from maybe some discrete coding use cases, the demand is very clear. But the big question is whether demand is going to materialize at the same level as the extraordinary level of supply we're building.
Timothée Lacroix [29:30] Yeah. Around this, I think the expectation is that trust, demand, and basically amount of tokens generated for the enterprise will completely jump once you are not bound anymore by humans asking questions or reading them. As soon as you have enough trust to have agents running in the background, as soon as you've set them to run a bunch of ETLs, as you've got them running lots of workloads and you've got them consolidating data and knowledge across your entire company, then you're not really limited by the number of tokens that humans can create or read.
Timothée Lacroix [30:09] And so I think everyone in the industry expects the demand to jump at that point. And the reality is, for this to happen, you just need a lot of boring software and control and things like this.
Matt Turck [30:14] It's amazing how much all of this is engineering, right, versus just sheer performance of models.
Timothée Lacroix [30:21] Yeah, it's a lot of plumbing. And the goal is to make all of this plumbing easy and easier and to make it faster.
Matt Turck [30:23] All right. And you said we were about a year away.
Timothée Lacroix [30:26] I'm not the most optimistic person. It might be faster.
Beyond Chat: The "Banger" Sovereign Use Cases
Matt Turck [30:41] Who knows? And we talked about use cases a bit already, but let's just put that one to bed because it's such an important question. What do you think are the kind of banger use cases in the enterprise? Let's assume all agents work in a workflow kind of way that you described. Based on either your industry watch or, more specifically, talking to your customers, what is it that is going to generate an amazing ROI beyond coding, which is pretty established at this stage?
Timothée Lacroix [31:33] Yeah, there are several dimensions to this. Coding is an obvious one. And to me, to get the full ROI of coding, you need customization because a lot of ROI is unlocked on sprawling codebases that are completely impossible to know for something that's been trained on the web. If you've got an enterprise that's been building its own domain-specific languages for years, you'll need some customization for an agent to come in and be competent in that respect. So coding is definitely a big one.
Timothée Lacroix [32:08] If everything comes true as I hope, I think there is still a huge jump in how we accelerate knowledge workers. And I believe the magical experience of: you go to your chat assistant, it's connected to your systems, and you can ask it anything about the enterprise, just hasn't been realized yet. And it's really obvious when you see the kind of queries that people are making, expecting them to just work. And to me, who's building the system, it feels like magic. Like, if you need to somehow send an email to three people and coordinate a meeting and also gather data from some BI system, it's just something that requires a lot more plumbing and capabilities than we have today.
Timothée Lacroix [32:45] So that's going to be a huge lift. And I think the last one, which is maybe closer to my heart, is really when we start to customize models to a kind of data that is particular to an industry. So typically, if we work in oil and gas, they will have seismic data that we can help understand and make sense of. If we work with computer-aided design, they might have full databases of specific data formats that are not widely understood by the most general models yet.
Timothée Lacroix [33:25] And if we manage to build a system where, with a light touch from us, in my dream world, we don't really have to intervene. It's all self-serve for the customers. They can consolidate that data and then build themselves a model that really understands what their actual private IP is made of and makes sense of this, then I'll be super happy. And I think there is huge value to unlock there.
Matt Turck [33:29] Great. Where does the edge fit in all of this?
Timothée Lacroix [33:56] There are a few reasons to go to the edge. First, there are some regions where it's more convenient to be able to work without internet, and there are also a lot of capabilities that don't necessarily require a huge model. So if you just need something that goes voice to action on any device, today, with typically the Voxtral models that we develop, this is doable. Again, an area where the more focused your use case is, the smaller you can make the model through fine-tuning or through just distillation in an even smaller architecture.
Timothée Lacroix [34:41] I think voice-to-action is going to be a big use case. I think it will simplify the current stacks a lot for these types of things. There are also some privacy considerations where you could imagine all of the context consolidation stays on your personal device. And for most things, you can deal with a small model that answers a lot of your questions. And then you potentially can gate what goes out to another cloud-based model. I myself take the train a lot.
Timothée Lacroix [34:53] I like having coding assistance. Having Devstral run on my laptop while I code on the train is comfortable, despite the bad Wi-Fi.
Matt Turck [35:10] And presumably there are some defense use cases as well. So you guys do quite a bit of defense work, as I understand it, with France, with Germany. I think you mentioned some partnership with Helsing. Is AI on drones and that kind of stuff a reality?
Timothée Lacroix [35:43] A reality? It's something that we work on. Yes, we have a robotics division that works with these partners. Having very well-defined use cases makes us able to really take the model down to lighter sizes. And it's, of course, use cases where control is super critical and you need to be able to really validate the solution.
Mistral 3 Architecture: Mixture of Experts vs. Dense
Matt Turck [36:16] All right, let's switch to the model part of the discussion. In December, you guys released Mistral 3, which was a big release, still with the MoE architecture, which is at the core of what you guys have been doing. You mentioned efficiency earlier in the conversation. Maybe walk us through the general thinking and approach in a highly competitive world of AI models, both in terms of closed source, but also very much open source and all the Chinese labs. What is it that you guys are trying to do, and how do you position?
Timothée Lacroix [36:56] Yeah, so we've released Mistral Large 3, which is an MoE. MoEs are really nice systems to train because of the lower amount of FLOPS, which makes us able to push performance a lot more during training. They are not necessarily the best format for on-prem deployment because, as of today, if you want to get the best efficiency out of a mixture-of-experts model, you require a lot of volume because you're looking at deployments across dozens of GPUs, usually. And to justify that amount of GPUs, you need to have the right throughput.
Timothée Lacroix [37:40] We are training large MoEs to get the best performance with the most efficiency during training. We're also continuing to train dense models at other scales because, depending on the environments in which our clients want to deploy, this might be the more cost-efficient solution. I think both architectures are still valuable on edge as well. Sometimes you just don't have the RAM capacity to deploy something like a sparse mixture of experts, and so going dense is helpful there as well. But yeah, definitely for training, mixture of experts and their lower FLOPS are very interesting.
Matt Turck [38:13] What is the ultimate goal of the model effort? I mean, clearly you guys are a frontier AI lab, but are you trying to create the best models and solve AGI, or are you trying to be the best open-source model compared to the Chinese labs or whatever open source eventually comes out of the U.S.? What is it that you're trying to do?
Timothée Lacroix [38:48] We're trying to get the best models that we can, and the models that are most useful for the use cases that we cover in enterprise. And so, typically, with the rise of agentic behavior, one thing that's very important is how you deal with various contexts, how you deal with various documents being added to the input. And so having the capabilities to do architecture iterations, really trying new things in terms of model training, is critical. So we're pushing the boundaries of what the current models can do with the compute capacity that we have, but we're also trying to focus on the things that are most annoying in our deployments today.
Timothée Lacroix [39:38] And so one of the considerations that has been solved with a few harness tricks is the context of those agentic systems. So it's visible typically in vibe coding, but it's definitely applicable to a lot of other use cases where, through all of the tool calls, you'll have to consolidate and summarize the context to be able to fit everything and have the model focus on the right parts. To me, this is just an artifact of the current architectures. We're trying to fit things in linear context windows where, essentially, the questions that we're asking aren't really necessarily all linear.
Timothée Lacroix [40:23] And so we rely today on the file system for this. And I think that was the big change and realization through vibe coding, is that agents are good enough at manipulating file systems that they can use this as a replacement for their context window, basically. They can select parts of what they want to read. They can select parts of the tool results. And this minimizes the context-length requirements. This is the state today. I think we can do much better.
Timothée Lacroix [40:32] And I think there are a lot of improvements to be done on those types of questions.
Matt Turck [40:34] Do your agents run on sandboxes?
Timothée Lacroix [41:00] It depends on the types of agents, but the answer would be yes. If it's coding agents, usually we have sandboxes that will let the agent iterate and run. I think the depth of the isolation will depend on the use case. Typically, if the file system is just representing textual context and you're not expecting the agent to do much action on it, then you don't really need a full sandbox. You just need some representation of that context as a file system, and it can be any sort of abstraction.
Timothée Lacroix [41:15] But if you are, I don't know, typically running asynchronous code development, then yes, you need a sandbox.
Matt Turck [41:37] Great. What is the current constraint that you guys are facing to make Mistral 4, when it eventually comes out, do much better than Mistral 3? Is that a question of more compute, or is that a question of data? And in particular, are you guys doing anything around synthetic data that you can talk about?
Timothée Lacroix [42:06] Definitely compute, and the current deployment that we have will help, as it's going to be giving us a lot more Grace Blackwell capacity than we had in the past. And so that's something that we're very excited about. And when you add compute, you also have to add data. And so we've been hard at work making sure that our data mixtures are as high quality as ever and growing in size. But as you mentioned, one of the ways to do this is through synthetic data.
Timothée Lacroix [42:41] In terms of where we use synthetic data the most, I think a lot of the interesting work that's happening is for the post-training part, where we can build environments that look similar to an enterprise and then try to synthetically create queries that are hard and that will require multiple hops. And so all of this work, in addition to the coding work, the reasoning work, is really what makes the final model able to perform in the various environments that we work in.
Synthetic Data & The Post-Training Bottleneck
Timothée Lacroix [43:26] So before, it was about acquiring world knowledge, and the web helps a lot with this. Now it's more and more about acquiring know-how, and for this, it's really about trying to find what our customers are trying to do, trying to replicate it inside of our training environments. And you mentioned post-training, and that's one of the key topics of the last 12 months in particular: this evolution of LLMs into systems with both pre-training and post-training and a lot of reinforcement learning.
Matt Turck [43:38] Where do you guys fall in that spectrum? Are you pushing a lot of reinforcement learning? Do you believe that pre-training still has room to grow? How do you think about it?
Timothée Lacroix [44:12] Yeah, so we're definitely pushing a lot of reinforcement learning. Everything still has room to grow. What I'm interested in as the CTO is really how you make all of the steps of the pipeline work well together and how everyone can develop most efficiently. Typically, what happens in post-training is that you will have a team that's working on improving code. You will have another team that's improving different enterprise behaviors. You will have another team that's improving instruction following. And so all of this, at some point, has to come together because customers aren't happy if you require them to deploy five different models to get their job done.
Timothée Lacroix [44:51] There is really an internal engine and capability around making all of these workstreams come together in the way that you expect that is super interesting to build. Internally, we're building and improving all of the parts of the stack. I think post-training is very rich because it also touches all of the new use cases of LLMs. And I think it's been very exciting to see just all of the new use cases that pop up every day. Anytime someone on Twitter finds new exciting things that they've done, then suddenly, you've got to make this proof of concept into potentially a base capability on which your model will perform well.
Reasoning Models: Why "Thinking" is Just Tool Use
Timothée Lacroix [45:12] And that's potentially an entire stream of work. And you've got to do this efficiently and prioritize well.
Matt Turck [45:23] Where does reasoning fall in all of this? You guys launched a reasoning model called Magistral a few months ago. Is that a big priority?
Timothée Lacroix [45:50] So reasoning is a big priority. And the interesting thing about reasoning was really how you can train models with reinforcement. And so it was first shown through reasoning because the system would learn to create better reasoning traces to get to better results. But the system is the same whether you create reasoning traces or whether you iterate on the tools that you call, or mix both. And so I think more and more the way to train all of this is going to come together.
Launching DevStral 2 and the Vibe CLI
Timothée Lacroix [46:22] And sometimes you'll have reasoning traces, sometimes they'll be long, sometimes they'll be short, sometimes there won't be any because it's not necessary. And there's no real difference between creating a new thinking trace or calling the right tool. It's all the same to me because what you're optimizing at the end is what is the best output for the model to create before it gets results to me.
Matt Turck [46:36] Great. Let's talk about DevStral 2 and the Vibe CLI. So walk us through those products and what they do and why people should use them.
Timothée Lacroix [46:55] Sure. So DevStral is our agentic coding model. And so it's something that you typically vibe code with, and you are more than welcome to vibe code with it through our CLI, aptly named Vibe. The value of vibe coding and why we focus on it: coding is a huge use case in enterprise, and especially a lot of our clients have large codebases where it's helpful for us to take our system and customize it to their codebase to let our agent run.
Timothée Lacroix [47:42] Now, DevStral and agentic coding is not only about vibe coding. The same system, when you run it asynchronously, can be used to review PRs. It can be used to check code for specific conditions. It can be used to modernize code. So its applications, even in coding, are quite wide. As I alluded to as well, having a system that is good at handling a file system is more generally very interesting. Even if you're not using it to code, you can use it to reason about enterprise knowledge, you can use it to connect to enterprise systems.
Timothée Lacroix [48:14] And to me, it's the basis of really the enterprise intelligence that we're starting to build. And so the big news is that those systems are going GA. We've got an offer where chat users, so Le Chat, our assistant, will also get the ability to use Vibe and the associated models.
Matt Turck [48:30] Another thing that you released reasonably recently, I believe, is OCR 3. What does that do? That enables you to just scan any form, any document?
Timothée Lacroix [48:55] Yeah, OCR is a huge use case in enterprise. A lot of our customers have—I mean, the typical example is KYC, where someone will submit a form and you need to input that information in a structured way in your systems, or you need to reason about it. And so OCR, interestingly, is not the type of system that I would have expected LLMs to really make large strides on. The visual reasoning and the visual understanding has gotten so good that it's just an easier way to process things in my mind.
Timothée Lacroix [49:22] You have any sort of input and you can get the data that you care about. As I mentioned, when you build agents, you have different types of inputs for the task that you're trying to solve. Documents and visual information are just a very, very frequent kind of input. Sometimes it's a lot cheaper to use a small OCR model to just get the text that you care about and then potentially post-process it or deal with it with another system than to run it through a large multimodal model that will basically do the same thing but at a higher cost.
Matt Turck [49:54] Yeah, you mentioned multimodal. To which extent is Mistral multimodal, or to which extent is that voice? Is video something that you guys either do or think about, or is that just not a big enterprise use case?
Timothée Lacroix [50:23] So to answer on the first part of the question on whether we build multimodal models, yes. It's always a balance between exploring in a direction, getting good capabilities and getting the first model out there, and then integrating it into the trunk, like the main model that we use for everything else. So those will always happen at separate times. But for audio, we have Voxtral, as I mentioned, and all of our main models understand images and can reason about them. For videos, it's a subject that we tackle through the lens of robotics first.
Timothée Lacroix [50:32] And so we're doing our first explorations on that topic.
Engineering Lessons: How to Build Frontier AI Efficiently
Matt Turck [51:11] Okay. Well, again, the velocity has been super interesting to watch. I, again, appreciate you reminding us that you guys have been doing this for only a couple of years. So just very impressive altogether. Maybe taking a step back and thinking of all this in terms of engineering and lessons for builders: as we alluded to a couple of times through the conversation, you guys are doing a lot with comparatively—it's very relative in the world of AI—fewer resources. How have you been able to do this from an efficiency standpoint?
Timothée Lacroix [51:46] We focused on the parts that we knew would provide the most impact, and we focused on basically what we could afford at different times. So when we started and we had enough resources to train a few models, we focused on getting the data perfect because we knew this was potentially not the most exciting part of the work, but it was absolutely critical. And any improvements in the data quality would 10x the improvements that we would get by really improving on the model architecture or things like this.
Timothée Lacroix [52:05] And so I think it's focusing the right effort depending on the scale of the company.
Matt Turck [52:30] And from a team-building perspective, how have you gone about it? The three of you, the three co-founders, have a deep background in AI. Are you these days focused mostly on building an FDE team, or are you still building this large research lab effort? And how do you think about the right ratio?
Timothée Lacroix [53:02] We are growing all of our teams: research, FDEs, product engineering, infrastructure for compute. And all of the teams have their own challenges in how you build and in what order you recruit people. It's been important to me and Guillaume and Arthur at the start. The three of us were good AI practitioners, so we knew how to train models and we knew how to code. And so we started with people like us to get the models trained the fastest.
Timothée Lacroix [53:31] But that doesn't work as you scale. It is critical to build the right infrastructure for research, and so this takes different skill sets. And it's something that we've been building over the years as well. And it's fascinating as someone who used to do research at a smaller scale to see the kind of systems that are involved and the gains that you can have at scale. In terms of engineering, it's kind of the same story, really, where you start with a team that's broad in its knowledge and self-sufficient and can iterate fast.
Timothée Lacroix [54:12] And then more and more, you bring in experts or people that have seen larger scale and will tell you, like, well, this won't work in six months, and so we should fix that now. So it's been super interesting growing the company and seeing all of the successive things that break at each scale and overcoming them through either changing the system, changing the organization, or building new things.
Matt Turck [54:28] How have you navigated the whole Europe-to-U.S.-and-rest-of-the-world dimension of this? I mean, you're very much the pride of France, the pride of Europe as well. Equally, this is a global race. How have you made it work?
Timothée Lacroix [55:03] So we work on all three continents. We have offices in Palo Alto, we have offices in Singapore as well. Most of our employees work from Paris. It's a good representation of what we're trying to build, which is a solution that's independent and that people control. And in this target, it doesn't really matter where we're from or who we're building for. We provide the tools, and the customer, the end customer, then owns everything that's built on it. And so I think it hasn't really been something that I've spent much thought on.
Matt Turck [55:16] So what should we expect from Mistral over the next couple of years?
Timothée Lacroix [55:47] Over the next couple of years, I would say diminishing doubts on the ROI of AI, ideally. So faster time to success, larger and larger use cases being built, and really democratization of building tools with AI in enterprise. I think this is really what I target for our customers. It should be easy, and most people should be able to accelerate themselves through the use of AI. I think we've seen this happen quite impressively for coding, and it should be something that happens a lot more widely.
Timothée’s View on AGI & The Future of Intelligence
Matt Turck [56:26] I was struck throughout this conversation by how pragmatic you are and focused on precise goals around enterprise success. What do you make of the whole rush to AGI conversation and people being AGI-pilled in San Francisco and other places? Is that something that you see happening, or does that to some extent not matter from your perspective?
Timothée Lacroix [56:58] It matters because the better your systems are, the more impressive things you'll be able to do, and it'll become easier and easier. Requirements I see for control and governance in enterprise make me think that even if I had some AGI-esque model on my servers right now, if I were to go into a large bank and say, "Here is a thing, please let it control everything for you," they wouldn't be happy to let it do it. And so I think building the infrastructure properly is quite key to following the progress of these models and really being able to quickly unleash all of their capabilities.
Timothée Lacroix [57:41] So to me, it's two directions that are necessary. You need to improve the capabilities of the model, and it's super exciting to do so. But the journey of making it trivial and easy for everyone to unleash those models on your enterprise workflows without really wondering what's going to happen is equally important and honestly super fun as well to develop. There are lots of super interesting questions.
Matt Turck [57:58] Wonderful. Well, Timothée, thank you so much for doing this deep dive on Mistral with us. It's been fascinating. Congratulations on everything that you've built again in this very short period of time, and excited for what's coming next. So thank you for spending time with us.
Timothée Lacroix [57:59] Thanks. It was a pleasure.
Matt Turck [58:20] Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already, or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.