“OpenAI’s Model Hacked Us” - Hugging Face’s Thomas Wolf

The MAD Podcast with Matt Turck · with Thomas Wolf, Co-founder and Chief Science Officer, Hugging Face

Thomas Wolf is the Co-founder and Chief Science Officer at Hugging Face. We cover why an AI agent testing cyber challenges pursued Hugging Face’s CyberBench datasets as a side quest, why closed models refused to assist during the incident while an open model helped analyze it, and why safety depends more on alignment than whether models are open or closed.

Watch on YouTube

Chapters

  1. 1:00 — 17,000 Attacker Events—and a Strange Target
  2. 4:28 — The Attack Was a “Side Quest”
  3. 6:13 — AI Training Runs Left Notes for Each Other
  4. 7:09 — Closed AI Refused to Help
  5. 9:47 — Fighting Back With an Open-Source Model
  6. 13:15 — Open vs. Closed Is the Wrong Safety Debate
  7. 15:46 — AI Agents Start Social-Engineering Humans
  8. 22:24 — The Three Walls: Sandboxes, Guardrails, Alignment
  9. 24:34 — “Neuralese”: Can Humans Still Read AI Reasoning?
  10. 25:28 — Why Monitoring AI Agents Gets So Hard
  11. 28:10 — Reward Hacking and the “Paperclip Problem”
  12. 32:02 — The State of Open-Source AI in 2026
  13. 33:47 — Router Models and the Enterprise Shift to Open
  14. 37:01 — The Real Economics of Open Models
  15. 39:41 — Can Chinese AI Models Be Trusted?
  16. 41:37 — AI Sovereignty: Who Controls the Switch?
  17. 43:16 — Why Western Open-Source AI Matters
  18. 48:16 — Is AI Heading Toward an Oligopoly?
  19. 49:41 — The Race Toward Recursive Self-Improvement
  20. 51:54 — Why Thomas Signed the AI Slowdown Letter
  21. 55:14 — AI Slowdown—or Regulatory Capture?

Transcript

17,000 Attacker Events—and a Strange Target

Matt Turck [1:02] To take things in order, so the OpenAI hack.

Matt Turck [1:20] So I know some of it is still being unpacked. I think OpenAI was on stage at Black Hat in Las Vegas yesterday as well, talking about this. So what's the two-minute version of what happened for people that may have heard of it but may not have followed everything?

Thomas Wolf [1:47] Yeah, for sure. I mean, about three weeks ago, on July 11th, we started to have some strong hints that a hacker was trying to penetrate our infrastructure. So, for context, we are pretty visible in the AI world. The AI world being central now in the tech world, we're pretty central in the tech world. So we do have regular occurrences of people trying to hack into our platform. That's a common thing since, I would say, the past two years, and we've strongly upped our security team.

Thomas Wolf [2:22] We now have a serious team, so we're kind of used to this. But this one was different because, first, it was massively parallel and in a different way than just a typical hacker parallel thing, in that many tracks were explored in parallel. And also, there were some very strange things happening. I would say just two things that were quite strange. The first thing is we could not really make sense of what the hacker was trying to access. So usually, hackers try to get the same thing.

Thomas Wolf [2:45] They try to get passwords, they try to get credentials, they try to get credit cards, they try to get the type of thing that they can sell back, basically. And this hacker was really focusing on a specific part of our infrastructure, which is maybe slightly less protected as well, but which is around datasets. So, we have a part of our infrastructure that hosts millions of models, but people maybe know less about us that we also host hundreds of thousands of datasets.

Thomas Wolf [3:26] And some of them are also used for evaluation. And in this case, this specific hacker was really interested in all the datasets that were called CyberBench. And so it took us some time to really try to understand, and also was using different types of tools than the ones we are used to. I mean, nothing really like mythos level, nothing groundbreaking that we would be like superhuman, I don't know, like alien-type technology, but just a different type of approach. And so, in the course of trying to process, we quickly had, like this, and we explained that in the blog post, we had more than 10,000—we had like 15,000, 17,000 total, I think—events.

Thomas Wolf [4:04] And in the course of trying to process that and understand basically what was really the target of this attack, we both started to suspect that this was an AI agent and not just a human attacker. And also, we felt a little bit powerless, and that we can talk about later, in that we could not use basically our typical closed-source codebase or closed-source API to process this thing. But that's another topic. So we wrote a blog post, and we managed to stop the attack quite quickly.

The Attack Was a “Side Quest”

Thomas Wolf [4:33] We wrote, as always, we are fully transparent. We not only open source in spirit, but also in practice. So we quickly published a full recap, or at least a detailed blog post on the event. And then, about a week later, OpenAI contacted us and told us that this was most likely something that happened as part of one of their model development or evaluation, basically. Typically, a model that might be the coming wave of GPT-6 or Astra, we don't know exactly this.

Thomas Wolf [5:01] And so that's when I think the whole event took another turn, because what people quickly discovered is that the model was not at all tasked with attacking us, but decided to do that as a side quest of something else. And this something else being, basically, the model was asked to solve a cybersecurity challenge, or a cyberattack challenge in this case. And the idea is that we want to know, and that's totally fair, we want to know how capable this latest generation of models is.

Thomas Wolf [5:38] And typically, people want to test them on some dangerous task. And some of these dangerous tasks that we want to know how good they are on is cyberattack. And here, some of these challenges were actually internal, but the model decided that because the challenge was too hard. And in retrospect, some of the challenges in this specific challenge called CyberBench or Exploit Gym—I mean, there's a couple of names, but that's roughly the same thing—some of these challenges are maybe just not possible to do.

Thomas Wolf [6:02] So the model is just tasked with doing something that's not possible. It's an exploit. So you're given a vulnerability in one software, and the model is asked to see if it can exploit this vulnerability to get, basically, full machine access. And some of them are just not possible. So the model tried everything it could. At some point, it decided that maybe it could find the solution of the challenge somewhere and just download the solution, just submit the solution instead of trying to solve it itself.

AI Training Runs Left Notes for Each Other

Thomas Wolf [6:43] That's what we've learned. Since then, and just yesterday, we learned that this was maybe even much wider, which is this might be across several training steps, in particular, even several training runs. And some of the previous training runs—that was the most impressive, I think, learning we had at Black Hat yesterday—may have left some notes for future training runs, which is, I think, mind-blowing. Mind-blowing. But yeah, we can also talk a lot about that.

Closed AI Refused to Help

Thomas Wolf [7:09] But we've been working on agent collaboration at Hugging Face recently as well on the science side. And we saw how good these agents are and how actually, I would say, tempted or driven toward collaboration they are. So I'm not so surprised by that, but I'm quite surprised that there was this message board internally that just stayed unnoticed.

Matt Turck [7:40] So to unpack some of this, you alluded to the fact that, to be able to defend yourself, the closed-source models were not available. And that's the sentence you tweeted, and that's one of the key aspects of this, which is so fascinating: the thing you said, the first autonomous AI attack was carried out by a closed model and defended against with an open one, which is basically the reverse of what everybody thought. So can you unpack that? What did you guys do?

Matt Turck [7:46] How did you go about it? And what does that mean for open source?

Thomas Wolf [8:08] Yeah, I mean, so what happened in the—so we have a couple of traditional cybersecurity protections, like Wiz or Amazon. We use a range of them, but we also have a stack, like many people, that is mostly based around Claude Code right now, which we use for many things. We use that for deploying, we use that for coding, but we use that also for operating and processing. And in this case, it's not only that Claude told us, "I'm not allowed to touch cybersecurity," but also Opus, which was the fallback, was saying, "No, I'm also not touching this thing."

Thomas Wolf [8:44] So basically, Anthropic was just saying, "We won't process anything about that, but you're welcome to apply to our cybersecurity program," with a link to an application form. But the thing you have to realize there, and I was mentioning also earlier, is when somebody is penetrating your infrastructure, they start to what we call move laterally, which is usually you have an entry point, but the destination is quite far. So they kind of find a way to compromise some of the credentials there to get progressively more access to your infrastructure.

Thomas Wolf [9:08] You kind of have to move fast. Like, it's a matter of at least hours and even more minutes so that you can stop them as soon as you can, so that basically the access and the blast radius stay localized. So you don't have time to apply for cybersecurity programs. It's not the moment you want to fill in, like, a Google Form or something and just have someone take time to vet if you're supposed to be given access or if it's not, or if it's too dangerous, and maybe interview you.

Thomas Wolf [9:42] That's just definitely not the way this is going to work. And I think in the future world, cybersecurity is going to be a big topic, and I think it will keep being a bigger, more important topic. It's a little bit naive, I think, just to think that every company is going to be part of the same vetted cybersecurity program by just one of the two big labs. I think it's a little bit crazy to think that you're going to have, I don't know, 100,000 vetted companies that progressively apply.

Fighting Back With an Open-Source Model

Thomas Wolf [10:07] So anyway, in this case, we said, well, we had to stop this now. So we basically tried all the open-source models that we had, and GLM, which is close to the state of the art right now. It was just before Kimi K2 that this happened. Kimi K2 is actually really good as well. It was just very good to process this. And basically, we could extract some of the patterns, and we could understand basically the hacker here was trying to access mostly the dataset.

Thomas Wolf [10:41] So we just rebooted this part of our infrastructure. We have a very simple, very flexible way to respawn pods and nodes. So this was how we ultimately stopped it. But I think it was very ironic because I think one year ago, roughly around the summer, most of the discussion about open source was this very simple mapping where open source was equal to unsafe and closed source was equal to safe. And that seemed very obvious in everyone's mind.

Thomas Wolf [11:04] And there was this idea that if we only have closed source, we'll be just fully safe. And if we only had open source, we'll be very unsafe. Well, everything that's been happening in the past month has been, I think, basically contradicting this very simple mapping. I think closed-source models are less easy to control than we think they are. On the other hand, open-source models, for some reason—it might change in the future—but currently are not trained so much.

Thomas Wolf [11:34] On actually, I would say, bad behaviors like that. So they're pretty bad at cyberattacks or deceptiveness, if you look at that. So it's a little bit hard to understand exactly where this comes from, in part because open-weights models tend to come with a very extensive technical report that explains how they are trained. With closed-source models, we can only try to guess. So it's quite funny as well, because I was seeing a lot of people trying to understand why GPT-5 was behaving like that, and they were using Kimi K2 as an example of how this should be trained.

Thomas Wolf [12:14] So they use this supposedly open-source, unsafe, very bad, dangerous thing to try to understand why the good thing that we don't know anything about is being trained. But yeah, that's how the world is right now. So I think it's interesting. More generally, I think in the future—and to be honest, I would say I'm cautiously optimistic around that—and I'm not specifically against closed-source models or ultimately pro-open-source models. I just think both of them are necessary. Just like we like to have closed- and open-source software, we'd like to have—I mean, I'm happy to run on a Mac right now, which is kind of a mix of both.

Thomas Wolf [12:47] It's based on a Unix kernel that was open source, but then there are components of it that are closed source. And that's great because I'm also very happy I'm not on Ubuntu right now. It's super easy to record this podcast with you for this reason. On Ubuntu, I spent a lot of time, like many of us, just connecting a microphone or whatever and trying to watch a movie with my girlfriend. My girlfriend was like, "When are we going to watch the movie?"

Thomas Wolf [13:12] I was like, "I'm almost there. I'm almost there. Still just installing the code or whatever." So I think both of them have advantages and drawbacks. And I think the world where the frontier is closed and there are not-too-far open-source models that you can use as well for many things is actually a pretty good middle-ground solution.

Open vs. Closed Is the Wrong Safety Debate

Matt Turck [13:36] To make sure I got it right, so what you're saying is that, to some extent, it's open source versus closed source, but it's more the state of the world as of right now, like the way the current closed-source models are designed and the current open-source models are designed, versus anything that's intrinsic to one or the other. It so happens that the open-source models right now are designed in a way where their guidelines or alignment philosophy allows them to be more reactive to a cyberattack.

Matt Turck [13:48] Is that correct?

Thomas Wolf [14:10] Yeah, I think for many aspects, the closed-open distinction is almost orthogonal to safe and unsafe. People don't understand that easily because it's easier to do bad mapping than to try to understand the subtlety. But that's the case. You can have very safe things in open-source models. You can have very dangerous things. You have different balances of safety and dangerousness. I mean, to take one example, last year, at some point, a lot of the discussion was around fake news and writing fake articles.

Thomas Wolf [14:40] That used to be a big, big misuse. That was the main one people were talking about, right? Today, there is, of course, a lot of fake news. There's a lot of AI slop. We even have a new word for that, right? It's even hard to find fully human-written articles. All of that is, or maybe not all, let's say 90%, to be fair, made by closed-source models, right? And there was a time where we were like, oh, if we have open-source models, everyone's going to generate articles everywhere.

Thomas Wolf [15:12] We could not control these articles. We could not control people creating newspapers. Well, the reality is that this was a very wrong view of a danger that would be specific to open-source models. There was a wider danger around sources of truth on the web. I think that's just one example, but I think the same is true. I think both closed-source models and open-source models should be more aligned. Right now, closed-source models are deceiving people.

Thomas Wolf [15:33] I think this is a huge problem, and this kind of question of how well we are able to align them toward not doing things that obviously are wrong. And I think, honestly, lies and all this should be obviously things that models should not be used for. But this might happen as well for open-source models in a few months. And so I think this risk axis is just kind of orthogonal to the question of whether it's open or whether it's closed.

AI Agents Start Social-Engineering Humans

Thomas Wolf [15:46] We should just find a way to solve that for both closed-source and open-source models.

Matt Turck [16:09] Okay, great. And to make sure we cover it, since we alluded to it at the beginning of this conversation, the AISI incident, which just happened and that you mentioned hit close to home. Can you unpack for us what that was, in a couple of minutes, and why that's a big deal?

Thomas Wolf [16:29] So this happened in an evaluation of the latest generation of models, GPT-5 and Claude 4, by an independent institute in the UK called the AI Safety Institute, so AISI, which has been there for quite some time. I think they are pretty competent. They were created for the first AI summit back in Bletchley. And what they do, basically, is they task models with solving some form of, say, cyberattack challenge, and they measure how good the models are at doing that.

Thomas Wolf [17:01] And some of this, again, might be possible or not. And so here the setup was a little bit, I would say, the setup was a little bit suboptimal in two aspects, which is the— but maybe let me first talk about what happened. Basically, the model was tasked to solve these attacks, like to attack and to penetrate this subnetwork. And what it decided to do, it decided to get one of the maintainers of a library that could be used to operate this Active Directory to merge malicious code, with the idea that if this malicious code was merged by this maintainer, then there would be an update at some point on the software.

Thomas Wolf [17:43] That was used in the subnet it was attacking, and then this would give it an entry point. And the way it did that was actually trying to social-engineer the maintainer into merging. So it created fake accounts, fake GitHub accounts, that came commenting on the pull request and said, "Oh yeah, you should really merge this. This is solving, like, a big problem I also have." And then when a human stepped up trying to say, "Oh, this looks actually like malicious code to me," it tried to kind of blackmail, almost, the human, only to say, "This is not important," or, "You didn't really understand."

Thomas Wolf [18:23] And then it actually tried to cover its traces, changing the past message. So there was a whole social-engineering aspect here, basically trying to get the maintainer to merge this malicious code. And this felt like a distinctive step up from just pure technical prowess. So you can attack a sandbox, for instance, by trying everything you can, which feels very mechanical and machine-like to me. But trying to attack the sandbox by blackmailing or mailing or social-engineering some of the maintainers, that's a very different level, I think, of thinking.

Thomas Wolf [19:01] And for me, of course, as an open-source maintainer, I've often been in this case where I have someone opening pull requests, and then people pile up commenting on this pull request, and I try to really understand, what's this? I felt very like I could have been the target of this side quest of the model, basically. That was very interesting, or at least very scary. But to be fair, there were a couple of misconfigurations. I mean, some of them are by design.

Thomas Wolf [19:30] So when this team ran this type of evaluation, they deactivated the cybersecurity guardrails, of course. Otherwise, the model doesn't do anything. And there's basically three levels. So let me try to explain a little bit how you can prevent models from doing bad things. The first level is you put it in a sandbox, which is, it doesn't have access to the internet. And here, for some reason, they didn't want to do that. I think that might have been a mistake.

Thomas Wolf [20:00] And the idea in their mind was, we want to let the model have as much potential for inventiveness as possible, so we'll give it access to the full internet. The main reason this might not have been a good idea is that the model was prompted in a way that made it feel like it was operating in a simulation. So the model could have actually fairly thought that this was fine to do anything because this was, like, a fake environment. So this is not super clear, but yeah.

Thomas Wolf [20:34] And the second thing is, then you have some guardrails that basically analyze the prompt and say no or yes to some prompts. So here, obviously, you want to deactivate this one. Otherwise, you just can't evaluate anything because they will just say, "No, this is a cybersecurity challenge. We don't let the model do that." But there is another aspect that's roughly at this level as well, which is you can analyze the reasoning of the model. You can process the chain of reasoning and try to detect when something bad is happening.

Thomas Wolf [21:06] And here they didn't have something like that in place. I think the main reason is probably that, until recently and until the OpenAI Hugging Face attack, people had maybe a little bit of a limited understanding of how good this model might be, or how far—maybe more how far—they might go in terms of side quests on the trajectory of solving this challenge. So I think people were still a little bit naive in that. So I would expect that in the future they will have way more monitoring and sandboxes.

Thomas Wolf [21:36] But the third level there, really deep, is that the model, even with everything deactivated—guardrails, sandbox—in my opinion, should really be very reluctant to tell a lie to a human and to try to blackmail or deceive any human. I think this is just generally, in any case, behavior that's hard to find justified in any context. So there should be something deeply in the model that aligns it and that makes it say, "Oh, this is actually something I don't want to do."

Thomas Wolf [22:06] Just like we, to be honest, I have kids, and just like the thing I teach my kid, which is, you just shouldn't lie. That's not a good thing in any context. So yeah, that's the deep question. And maybe last year, I would say we would've thought that this was pretty good. And we had all this discussion around Claude's Constitution, model specification. And most of this model specification or constitution says you should be honest, you should not tell lies to a human, to any participant.

Thomas Wolf [22:23] And we thought that maybe this was kind of a solved problem. And what we see today is, it's not sure that's solved as much as we thought.

The Three Walls: Sandboxes, Guardrails, Alignment

Matt Turck [22:53] So, to play it back, you're saying there are those, what you call the three walls: there's the sandbox, there's guardrails, and then there's the model's alignment. And I think you said the sandbox and the guardrails only work as long as we humans are smarter than the AI, but that may only last so long. And therefore, ultimately, security is fundamentally an alignment problem.

Thomas Wolf [23:13] Yeah, I agree. And that's something, as you can understand, that's both the case for open-source and closed-source models. Ultimately, you want them to be aligned. Open-source models have the specificity that you may choose one, you may control where you want to run them. So it's harder to make sure everyone uses sandboxes and guardrails. I mean, we can definitely have some laws and regulation around how you should deploy these models, which we're going to have at some point.

Thomas Wolf [23:45] But it's a bit harder. But alignment is really the critical part, in my opinion. For the two others, I mean, sandboxing is what we've seen this year, and we've seen many examples: they are pretty much easy now for these models to escape from. It's really hard nowadays to say, "I'm going to make a fully foolproof sandbox. I'm sure it's going to be resistant against all the coming generations of models." I think we should assume that a sandbox will always have a small probability of not containing a model.

Thomas Wolf [24:12] But even beyond that, we can't air-gap the world. You can't just sandbox everything. Things have to talk with each other. We want our models to be able to do web search. We want them to be able to do stuff on the internet for us. So we can't just sandbox everything. And so what remains to us before alignment is just guardrails and monitoring. And I think these are, for some reason as well, as model capabilities become really good, also as the model—I'm a little bit worried that the model starts to talk in a form of English.

“Neuralese”: Can Humans Still Read AI Reasoning?

Thomas Wolf [24:34] I mean, I'm French, so maybe it's partly my problem, but I feel like they start to talk in a form of English that's harder and harder to process. That's very content-dense. They start—

Matt Turck [24:35] You call that neuralese?

Thomas Wolf [24:57] Yeah, it's not fully what we would call neuralese, but I think it's a little bit on the way to having a more and more difficult time fully understanding what the model is telling you. And it's not because the model is dumb. I think it's because probably part of it is because of the training process and how they are trying to be efficient, how they use their tokens. But this means that they start to bundle a lot of semantics in some tokens.

Thomas Wolf [25:07] And generally, it just seems like it's harder and harder for humans to fully understand what's happening.

Why Monitoring AI Agents Gets So Hard

Matt Turck [25:29] And just so that's in the chain-of-thought where the model explains what it's doing and the steps that it's going through, what you're saying is that it used to use perfect English, and now it's starting to use a different kind of language that you call neuralese, which is increasingly harder for humans to understand.

Thomas Wolf [25:51] Yeah. And I mean, it's just a very big simplification of all of that because there's a lot of research that says that, basically, you can't read everything in chain-of-thought. Not everything is explicitly said. But I would say, more generally, I think just relying on being able to read the reasoning trace to fully understand what's happening is also not fully bulletproof, in my opinion. And I take this neuralese as an example because I feel like a lot of people start to have a bit of this problem with Claude Code.

Thomas Wolf [26:22] So I feel like it's something that people can understand. But more generally, I think longer term, it's really hard to fully, fully rely on this only. And the same, I would say, is that you could say, maybe I just don't care about understanding exactly what's happening, and maybe I can just look at the tool call, and if I see a tool call that's bad, I can just block that. And I also think this is probably not bulletproof. And the way you can see that is probably three big things that are compounding.

Thomas Wolf [26:47] The first one is we start to use this model for many, many, many things. So as we talk right now, I have a model deploying a box somewhere on a VPS. It's doing many different calls. They're all going in different directions. I also have other models that I use for administrative tasks. So just the range of tools that these models are using is really, really large right now. So it's getting harder to say you're allowed to use that, but you're not allowed to use that.

Thomas Wolf [27:17] And this is the bad thing. It's getting very hard as we deploy them wider. They also work on larger and larger tasks where they use many, many things. So sometimes I ask it to do some coding, but the coding involves searching on the web and maybe doing these things, and actively using many things which are not just purely writing code and running some tests. And this test might be quite extensive and involve other software. So the frontier is much more blurry.

Thomas Wolf [27:27] And then you also have this swarm of multiple agents. So it's also harder to say it's all in one context. It might be split between many contexts. And maybe this subagent is doing something that looks pretty innocuous, but maybe combined with these other subagents, they're actually not so great. So there's all of these things that make it, I think, really harder to be fully sure that you have an exact idea of what every swarm of agents is doing.

Reward Hacking and the “Paperclip Problem”

Thomas Wolf [28:10] You probably have to really zoom out at the very global level and see what's happening in every direction. But that's a whole monitoring setup that we need to build right now. So, yeah, this to be said, I think as we deploy, how we use this modeling very complex, long-term parallel setup, I think it's going to be harder to just say, "I can look at the tools and I know if it's doing something great or not."

Matt Turck [28:26] Is there something fundamental to the way those models are currently trained, so the very frontier, that makes them more likely to go onto those side quests and potentially create harm? I mean, an analogy that people have been talking about for a very long time is the paperclip paradigm, which I think was Nick Bostrom in 2003 saying that AI may harm us not because it's trying to harm us, but just as a result of being given a goal and pursuing that goal relentlessly until it achieves the goal.

Matt Turck [28:55] So, are we in that world?

Thomas Wolf [29:21] And if so, what causes—yeah, that's a bit what I hinted at. I mean, it's always hard to be fully affirmative there for one reason, which is that we don't have full visibility on how the frontier models are trained right now. What we know, though, is we moved from this pure human data paradigm that was first just pre-training on human data and then also aligning with human preferences. That was called RLHF, where we had a lot of humans in the loop and human data.

Thomas Wolf [29:56] To a recent paradigm where models are trained a lot in this RLVR. So, basically full RL environments where they're allowed to explore and they just have one goal, which can be make this code pass this test, or can be capture the flag in cybersecurity, or can be install this. But this goal is a goal that's unrelated usually to any human preference or any moral or ethical or whatever deceptiveness. And it's a goal that's very just cold, true-or-false goal. And we moved to a paradigm where this is increasingly a very, very large part of model training.

Thomas Wolf [30:19] So this was this recent evolution. And that's also the paradigm where can happen what you were saying, which is you can have reward hacking, this type of thing, which is you actually solve the problem, but not using what was expected for you to use. So it can go from a pretty benign one, like what happened for OpenAI, for instance: I just tried to get the answer from somewhere I'm not allowed to, or to, like, a more harmful one where you actually have some impact on the human. It can be a GitHub maintainer for now, and later can be another type of humans.

Thomas Wolf [31:10] So it seems to be way harder to make sure that what we had kind of solved, or at least what we were doing pretty well on in the full human-driven paradigm, also apply in this kind of more machine-driven paradigm, I would say. But that's a hint, I think. But definitely, it seems like when Bostrom wrote about it in 2003, it seemed a little bit futuristic, definitely, and maybe something that was a little bit crazy and just would not happen. But today, I mean, it's pretty clearly something that happened, and it's the best description of what we've seen the past two weeks, this type of thing.

Thomas Wolf [31:42] GPT-5 and the o3 models don't seem to have at all the same type of behaviors. So there are differences here in the effect of—they are not trained exactly the same way and they don't behave the same way. So that's a pretty positive sign, in a way. That means that we can actually probably tweak this to go in the right direction. But the best way would be to know a little bit more about how they're trained or what they try and what doesn't work or what should work.

Thomas Wolf [31:58] I think that's kind of the idea of open science, and that's something we advocate a lot at Hugging Face.

The State of Open-Source AI in 2026

Matt Turck [32:35] Fascinating. And so, speaking of which, let's zoom out a bit. We'll go back to maybe some of the implications in terms of policy of all of this. But since you mentioned the ever-so-important role of open-source AI, what's your sort of quick, high-level take on where we are in terms of the state of open source? So there's been this race between open-source models and closed-source models. Depending on who you ask at what time, open source is about to catch up.

Matt Turck [32:53] Sometimes open source is just as good. Some people say no. What is your sort of realistic, pragmatic take on the current state of open-source AI?

Thomas Wolf [33:12] I think it's very strong. 2026 is maybe the year of cybersecurity, but that's also very clearly the year of open-source AI. I mean, first, all the doomers that were saying open source is not going to be able to stay close to the frontier, I think they're—at least up to now—they've been pretty wrong. I mean, it's also clear we don't have any Mythos-level open-source model, for sure, but we definitely have models that are not super far from the Opus category, or depending.

Router Models and the Enterprise Shift to Open

Thomas Wolf [33:49] Also, it's more and more spiky, so you need to find your spike. Some people stand on some spike or not, but typically they're definitely pretty good right now. And they've been following rather closely the frontier, at least on the benchmark. It's also not like it was maybe in the early days of benchmarking, like we say, when your model is only good on the benchmark, but it's very bad as soon as you leave the benchmark. A lot of these models are pretty generic in their good capabilities.

Thomas Wolf [34:20] So, yeah, it's very good. I think there are two strong trends I see right now. The first one is I see a move in companies to try to control their costs. So there's increasingly discussion there. Maybe 2025 was the year of token maxing, where you could say, hey, you should spend as much in tokens as you're paying your employee. This year, people realize that actually we spend a lot of money on salaries. So if we spend the same exact amount, that's going to be basically doubling our costs.

Thomas Wolf [34:46] So, yeah, which seems pretty obvious in retrospect, but that's quite true. And not every company can assume to double their costs. It's also pretty stupid right now to just say, we're going to fire everyone and work on agents. We all know they sometimes go not directly in the direction we want them to, and you need humans to shepherd them. So I think a lot of companies are trying to find what we saw a lot, which is kind of a fusion or router model where you use the frontier for something, but you find the smart way to gracefully fall back on less expensive models for simpler tasks when you don't need to.

Thomas Wolf [35:33] Even in our daily life right now, when you code with a frontier model, in many cases you ask it to spin out subagents, and they might use lower-performance models. It can be Claude using Terra Luna, it can be Claude Code using Opus, Sonnet, or Haiku. So I think everyone's getting, even at the frontier and the closed-source model world, used to employing different types of models. And then it's very natural that some of these could be really, like, very cost-effective agents. And most of the time, you want to go to open source in this case.

Thomas Wolf [36:05] There's a lot of cases as well for the strong ecosystem of inference providers. Fireworks has been on a roll. All the clouds, Nebius, CoreWeave, every cloud has increasingly had these crazy revenue curves that are just basically a translation of people using more open source. So, yeah, I think open source is having a very good time in terms of staying solidly close to the frontier and driving more adoption. And the second big trend, I would say, is it used to be only China, but there is also now a range of promising companies in the West.

Thomas Wolf [36:42] It used to be that only Meta was open-sourcing for everyone. I mean, Meta kind of left the field, but they might come back, who knows? But yeah, this void was filled rather quickly by companies like Reflection AI, Thinking Machines Lab, Arcee, Mistral is supposed to open source a model soon, NVIDIA themselves training models and actually training very good models right now. So, yeah, I think there is a range of very promising teams there that I'm quite bullish on, to be honest. And maybe between even the time we are recording and the time you actually release the podcast, we might see another couple of very nice Western models released.

The Real Economics of Open Models

Thomas Wolf [37:01] It would be great. Like, open source doesn't have to be synonymous with just one country making them. It could be, like, a global thing for sure.

Matt Turck [37:32] On the first point, on the sort of enterprise adoption of open-source AI, there was a moment in time when people associated open source to free, but I think the world has quickly learned that while the models may be free to download, deploying them and serving them is certainly not free. What is your sense of the cost advantage of open source in reality in the enterprise?

Thomas Wolf [38:00] Yeah, that's a very good note. It's the same very stupid mapping: open source equals free, which is also wrong there. And it's much more subtle, and, like, you have an advantage or not. I think the nice thing about open source is you have quite a wide ecosystem. Like, the entry barrier to be a cloud provider is pretty low. So you have a lot of competition there on what's going to be the cost per token. And this is a mix of many things, right?

Thomas Wolf [38:27] It is: how cheap are you renting your data center, or how cheap can you buy your chips? And here, actually, we have also new chip companies who are going to come on the market, and this is going to be very interesting to watch. And so, basically, how cheap is your hardware, and then how much can you optimize the model? Can you quantize them? Llama 2 was quantized by NVIDIA in 4-bit. And this is to make it faster and smaller to run.

Thomas Wolf [38:55] So you have a lot of strategies you can use and explore to make this model cheaper. But it's also true that they don't have to be cheap per se, and also that the closed-source model may be, in a way, subsidized right now. Like, the amount, the number of tokens you get for your $20 ChatGPT or Claude subscription might not be the full price that they actually pay for your tokens. So there is this kind of complex balance around costs, while maybe your cloud inference provider for open-source models doesn't have so much leverage that you can lose money on the subscription business.

Thomas Wolf [39:12] So we'll have something more complex, I think, ultimately—

Matt Turck [39:18] All subsidized by venture capitalists.

Thomas Wolf [39:27] Exactly. We have the same thing we've seen in many fields, right? But we also know this is ultimately a little bit temporary. So you should not fully rely on that.

Matt Turck [39:35] This is not the equilibrium. Yeah, sort of the Uber phenomenon, right? Like, cheap Ubers before the IPO and expensive Ubers since.

Can Chinese AI Models Be Trusted?

Thomas Wolf [39:41] I think you want to keep the ecosystem alive for the post-VC market.

Matt Turck [40:14] On the second point, China versus Western models, does provenance actually matter? What is the latest thinking in terms of if your model is completely open and then you know exactly what's in it, then you shouldn't worry, it's completely safe? Is this still a little bit of thinking at the back of people's minds that there might be some backdoor, some trickery to Chinese models? Is that still a current question?

Thomas Wolf [40:41] Yeah, I would say that's a good question, of course. And the thing about open source: open source doesn't know any borders. You can't really keep your open-source model restricted to download to just a subpart of the Earth. So by default, it's kind of a global thing, and then you want to maybe understand the provenance and the supply chain. So there has been some work on this, the sleeper agents. I think it's probably under research. I think there is mostly work by Anthropic on that, and it should probably be reproduced, be explored deeper, to understand how much it's possible to implant kind of a backdoor that would be triggered by a prompt.

Thomas Wolf [41:19] Also, to be honest, even for the closed-source models right now, we have some struggle controlling them fully. So I don't think we're very clear on how we control even the closed-source models at the moment. So, yeah, I would say it seems to me that it could be a possibility for sure. We have not seen any indication of that. It's very easy to fine-tune them, and you can change quite a lot the weights as well. So right now, if you pre-train and post-train a model for longer, you very likely change quite a lot the weights that it has.

AI Sovereignty: Who Controls the Switch?

Thomas Wolf [41:50] So I think there's a lot of ways to circumvent that, which means that at the moment I'm a bit less worried about that than maybe just pure reward hacking that we actually have already. Yeah, I mean, just generally on sovereignty, I think sometimes people are a little bit confused there. I think the most important thing is who has the hand on the switch to trigger or not your intelligence access. So I think when people really realized that was earlier this year, where the U.S. decided that Fable was only accessible to U.S. citizens.

Thomas Wolf [42:27] I think at least in Europe and Asia, that's really the moment we saw governments understanding that someone had a trigger and they could say, no, you're just not allowed to use this intelligence anymore, like this token. I think that's the critical thing you should, at least for me, that's really the level one of sovereignty, which is: can someone just decide that you don't have access? Like someone, I mean, like a country-level thing. And this has two things. This has the API access, so it can be the data center.

Thomas Wolf [42:57] So if you don't have the data center, it means also, like, another country could say these data centers are not accessible anymore to this citizen, for instance. So I think that's really the first thing you should see. And then there's a lot of future questions, but you should build your stack keeping that in mind. So an open-source model, you can download it. Nobody can, like, no country could take it out from you once you download it. You can fine-tune it yourself.

Why Western Open-Source AI Matters

Thomas Wolf [43:16] If you operate it in your data center, like a local, on-prem data center, I feel like you start at the beginning of a sovereign stack. And then there is obviously a lot of questions around backdoors and more complex stuff, but that's kind of the basic minimal thing in 2026, I think.

Matt Turck [43:46] And as an open-source optimist, do you worry about the motivations for Western open source to thrive? So China has clearly a geopolitical motive behind being at the forefront of open source. But if you think of the West, then, so you mentioned NVIDIA, and we had Brian Catanzaro from the whole Nemotron effort on the podcast a few weeks ago. So clearly there is a motivation for NVIDIA to do this, which is they sell the chips, or having thriving open source all makes sense.

Matt Turck [44:24] But if you think of everybody else, it's sort of unclear why Western companies would do open source. Reflection might be the exception, but the models haven't come out as far as I know. So do you think about this motivation and what that means for the future of Western open source?

Thomas Wolf [44:48] Yeah, of course. But that seems pretty obvious to me, that we actually want that. And I think a lot more people have incentives than you may think. I think as soon as you're interested in having a thriving business ecosystem with many companies being able to actually use AI, and not just as a thin wrapper around another company, but as real AI builders, I think you want some open-source models. So that's one of the reasons the U.S. government just recently said, actually, we want to keep open source thriving.

Thomas Wolf [45:17] And basically the idea is that open source is one of the best ways to get new business. So it can be for many things. It can be just because it allows you to not just end up with an oligopoly, basically, of two companies, which, I mean, we've seen many cases of oligopoly in the past. It's not always the best thing for competitiveness, for price. There's many dangers with that. I think just rushing in the direction of saying, we've taken our winners and these two companies are going to build AI for everyone else.

Thomas Wolf [45:45] I think that's, from a business, economic, market-side point of view, I think, doesn't really seem optimal to me. And then it also limits a lot of invention. So just to take one example, there's a lot of potential right now in biology. There's a lot of new life science companies. There's a lot of companies wanting to explore that. Because of the guardrails and because of the question around biohacking and using these to generate, like, the access right now for LLMs, just to take it, is very, very limited once you want to ask some biology question.

Thomas Wolf [46:23] And so basically most of the life science companies I've seen who were using this model had to switch to another option if they wanted to be able to process anything related to biology. So if you're a new company exploring something that the big labs are not currently, at least, exploring, or they don't feel like there is enough business potential maybe to give full access to everyone, I don't know exactly, but basically you cannot really use that. So either we say all life science is going to be built from now on by Anthropic, OpenAI, and Google, maybe Meta, or we say we want a big ecosystem around there.

Thomas Wolf [46:58] I mean, my personal opinion, and I'm biased, is that you don't want too much concentration of power around key technology. And I feel like the more people can invent, the more we have a diversity of new ideas, diversity of new companies, but also as investors, right? And you're investors, you know what I talk about. Like, let's say you could only invest in Anthropic and OpenAI. That's a little bit sad, right? It's a little bit boring. It's only for growth stage, but you want to invest in new companies, and you don't want to invest only in this very thin wrapper.

Thomas Wolf [47:31] You want to invest in new companies that actually are able to build AI. And most of them, and we see them at Hugging Face, most of them need open-source models. Another big example is all the things around gaming, video, or even robotics. Most of the time, what they do is they start from an open-source model and then they fine-tune it on some robotics data. For gaming, for instance, they'll take a video generation model that's open source. They need to do something that was not predicted by the video generation startups.

Thomas Wolf [47:47] They need to add action. So what they do, they fine-tune with action in the loop and video, and that's how the first, I think, interesting, like, real-time gaming companies started, basically. So a lot of the time, open source is your easy way as a new company to start to have your own models, to fine-tune on your own data, to be able to start to build your own AI, and not just to basically sell your training data back to the model providers, which is, I think, always dangerous because a lot of these model providers might at some point want to enter your field if you're basically selling them your data.

Is AI Heading Toward an Oligopoly?

Thomas Wolf [48:16] And this happened in the past already in legal, in design, in many fields, I think.

Matt Turck [48:50] To play it back, so your take on the July 24th industry letter on open weights and American AI leadership, which was also Jensen's first tweet ever, that you guys signed, obviously, is partly important and partly a resistance to just an oligopoly structure that is being put in place. So it's both: it's good for the world, but there's a strong economic motivation behind it. Is that fair?

Thomas Wolf [49:16] Yeah, I think everyone believing in inventiveness and being able to create new things also in the AI world would want a part of open-source access. Just like, if every code was closed source, we would not have the thriving coding ecosystem we have right now, right? That's kind of obvious. Everyone would have to work at one of the large closed-source software companies if they wanted to create any software. That doesn't seem even really possible, to have all the inventiveness and creation we've seen in the software industry.

Thomas Wolf [49:31] I think the same happened in AI, and I don't want to dismiss risk, and I fully agree we need to work on alignment.

The Race Toward Recursive Self-Improvement

Matt Turck [49:57] And I would say, for now, open-source models being under the frontier, I think this is maybe less important than some people wanted to. Maybe to take a step back as we get near the end of this conversation and sort of get a sense for where the world might be going from your perspective. In your post yesterday about ASI, you talked about a new wave of labs. So, in particular, you referenced the news of Geoff Dean's new company that he just announced yesterday.

Matt Turck [50:31] I mean, yesterday was a bit of a crazy day in terms of everything that came out in the world and was announced. So, the point being that this company is explicitly rushing toward recursive self-improvement. So, in the context of everything we described, how nervous are you about this evolution towards self-maintaining, self-developing, recursive AI?

Thomas Wolf [51:03] Yeah, that's a good question. I'm definitely, as a scientist-researcher, I would say I'm very interested in the idea. I feel like there's a lot of these superintelligence labs that want to tackle solving crucial challenges for humanity. I think this is a great goal. I would love to see AI making more scientific discoveries. I think that would probably be the most beneficial thing that AI could bring, more than just AI slop everywhere. So I think I'm very optimistic. I just feel like, I would say, the past few weeks has raised a little bit the question of how good are we at aligning these models.

Thomas Wolf [51:37] So, like in French, we say we don't want to put the cart before the horse. We need to go in order there. So, yeah, I would say right now, the good thing is most of these seem to be internal research labs. Hopefully they do good security around what they do before they deploy some of their product. They think how it's going to be used. They have a good feeling around, I would say, good thinking around the social impact of what they're building.

Why Thomas Signed the AI Slowdown Letter

Thomas Wolf [51:54] But yeah, I still think that we should try to understand really well how we're going to deploy this model in the human world, I would say.

Matt Turck [52:22] Right. But you're not in the camp of the petition that came out, I think, four days after the Open Weights letter that NVIDIA did, that we were talking about a minute ago. There was a different letter that came out with 1,100 people, and this time both Anthropic and OpenAI signed, that asked government to help deliberately pace the frontier of automated AI research. So basically, the industry asking for a slowdown. Are you in that camp, or do you think that's just not the way it works?

Thomas Wolf [52:51] I did sign this letter. Yeah, I agree. I agree that I think we should go there. I'm also—and you probably have the same feeling as an investor—I have the feeling that even if we stop right now, we would still have quite good companies we could build on top of what we have right now. I feel like there's a lot of things we can already do with these models that are already extremely interesting. I feel like there's a lot of things we need to understand and should do in terms of open science.

Thomas Wolf [53:20] And sharing how they work. So I'm not in the camp of, we need to rush really quickly right now. The main question is, if we want to slow down a little bit, that would also be great because maybe then we don't have four announcements per day that we need to mix in one podcast. It's not the next podcast tomorrow. Maybe I can take one day of holiday in the summer. But no, I think the main question is, can we do it right?

Thomas Wolf [53:40] That's the main question here. I think a lot of people would be fine with AI going a little bit slower, being a little bit more open, being a little bit more caring, being a bit more reflexive, and trying to understand better how to do that really well. But the main question is, how can we negotiate and organize a slowdown without having bad incentives, where just one or two players not slowing down will kind of break the whole effect of having a slowdown?

Thomas Wolf [54:08] And so, yeah, I don't think the letter gives any incentives. It might need some collaboration. There was one long blog post called AI 2040. I don't know if you read it. That was also advocating for kind of a careful slowdown, and maybe somewhere between we go full brakes out, we go as fast as we can, and somewhere between we regulate everything so nobody uses AI, which I think are both stupid solutions, but something around, we try to see if there is a way we could actually pace this a little bit slower.

Thomas Wolf [54:59] I think that would be great. I don't think we would lose a lot. And I think actually, also in terms of company creation and all of that, we could still have a lot of really great things happening. But yeah, I'm actually sympathetic to both this and open source. I don't think open source has to be acceleration per se. Dan was saying this is decelerationist. I don't think it's also decelerationist or deflationist. I think these are also orthogonal.

AI Slowdown—or Regulatory Capture?

Thomas Wolf [55:14] You can be pro-open science, you can be pro-openness, and you can also think that actually we need to understand how to train this model well, and we need to actually be able to do real science right now.

Matt Turck [55:37] And you're not worried about this being an attempt at regulatory capture, where the top private labs are effectively trying to figure out how everybody else can slow down? I mean, you mentioned the risk of not everybody just complying, but effectively freezing the market structure around who's a leader and who's not.

Thomas Wolf [56:07] Yeah, I don't think it has to be. I feel like you definitely have the same path that's actually fully accelerationist, where you decide a couple of companies are racing against each other and you regulate all the other actors out. I don't think regulation has to be synonymous with slowness. And definitely, the question is more how you put that in action, how you actually put that in practice, and how you deploy this in regulation or cooperation. I think Demis also had a pretty nice letter the other day before he stepped down or up as chief scientist. Still don't really understand where he's going to be now.

Thomas Wolf [56:45] But his letter for basically international collaboration was also very much pro-open source in some aspects. I think you can have a slowdown that's very open source. That's the one I would love to see, which is we slow down and we use the fact that we slow down to be able to actually share more things. And I feel like a race dynamic is usually more in terms of closing the doors of the labs, right? So, to me, a slowdown is probably more the opportunity to open.

Thomas Wolf [57:02] But of course, if it turns out to be mostly a way to just solidify, like we were saying, a cartel or oligopoly of just two companies, I'm not very excited about this direction.

Matt Turck [57:18] Wonderful. Well, that feels like a wonderful place to leave it. Thomas, thank you so much. This was absolutely fantastic. Really enjoyed it, and appreciate you taking some time to speak with us in the middle of your time off. So thank you so much.

Thomas Wolf [57:21] Appreciate it. Thanks, Matt.

Matt Turck [57:41] Hi, it's Matt Turck again. Thanks for listening to this episode of The MAD Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing, if you haven't already, or leaving a positive review or comment on whichever platform you're watching or listening to this episode from. This really helps us build the podcast and get great guests. Thanks, and see you at the next episode.