Chasing Real AGI: Inside ARC Prize 2025 with Chollet & Knoop
The MAD Podcast with Matt Turck · with François Chollet, Co-founder, ARC Prize and Ndea
François Chollet is the Co-founder at ARC Prize and Ndea. We cover why ARC measures adaptation to novel problems rather than memorized skills, how test-time search drives major gains at far higher inference costs, and why ARC-AGI 2 tasks take humans about five minutes while base LLMs score 0%.
Chapters
- 2:07 — What is ARC and how it differs from other AI benchmarks
- 2:54 — Why current models struggle with fluid intelligence
- 3:52 — Shift from static LLMs to test-time adaptation
- 4:19 — What ARC measures vs. traditional benchmarks
- 7:52 — Limitations of brute-force scaling in LLMs
- 13:31 — Defining intelligence: adaptation and efficiency
- 16:19 — How O3 achieved a massive leap in ARC performance
- 20:35 — Speculation on O3's architecture and test-time search
- 22:48 — Program synthesis: what it is and why it matters
- 28:28 — Combining LLMs with search and synthesis techniques
- 34:57 — The ARC Prize structure: efficiency track, private vs. public
- 42:03 — Open source as a requirement for progress
- 44:59 — What's new in ARC-AGI 2 and human benchmark testing
- 48:14 — Capabilities ARC-AGI 2 is designed to test
- 49:21 — When will ARC-AGI 2 be saturated? AGI timelines
- 52:25 — Founding of NDEA and why now
- 54:19 — Vision beyond AGI: a factory for scientific advancement
- 56:40 — What NDEA is building and why it's different from LLM labs
- 58:32 — Hiring and remote-first culture at NDEA
- 59:52 — Closing thoughts and the future of AI research
Transcript
What is ARC and how it differs from other AI benchmarks
Matt Turck [2:11] François, the big news is the announcement of ARC Prize 2025 and the concurrent launch of ARC-AGI-2. But maybe to start from the top, what is ARC-AGI as a benchmark, and why did you create it?
François Chollet [2:46] ARC is an AI benchmark that tries to measure AI fluid intelligence as opposed to skill at specific tasks. And that's a very different approach from basically any other AI benchmark. Most benchmarks, they're looking at the ability to answer specific questions, perform specific tasks. That's going to be very reliant on knowledge or skills that can be memorized in advance of trying to pass the benchmark. And ARC is not like this at all. ARC is a set of tasks that you cannot prepare for. So it's not trying to measure what you know; it's trying to measure how well you can adapt to something you've never seen before on the fly.
Why current models struggle with fluid intelligence
François Chollet [3:22] As it turns out, this is actually something that's extremely challenging for AI. Current models, they're really good at absorbing mountains of knowledge and specialized skills, but they're not very good at making sense of novelty on the fly, at adapting. This is something that AI is not good at, and that's what makes ARC really interesting. And one thing to note is, last year there was this big shift in the AI research world where the AI research community started to move away from this idea of just scaling up pre-training and then using the pre-trained models in a static fashion at inference time, just running them.
Shift from static LLMs to test-time adaptation
François Chollet [4:15] They're not changing, they're not learning anything new, and they're not adapting to novelty. GPT-5, Gemini, that kind of thing. And there was a shift towards instead doing test-time adaptation, in particular test-time search, which is what you see in a model like, for instance, o3 or o1 Pro. So models that have much, much longer latency, much higher latency at test time, that are also much, much more expensive. We're talking 10x, 100x more expensive, and that have much stronger generalization power. And there was really only one benchmark that provided a strong signal about the fact that this massive shift was happening.
What ARC measures vs. traditional benchmarks
François Chollet [4:53] And it was ARC, because all the other benchmarks, they were really looking at skill. So they were not able to differentiate between just brute-force scaling and memorizing more and more knowledge and more skills versus actual IQ increases. And ARC is really the only benchmark that's looking at effectively machine IQ. And that's what makes it a really interesting signal. So why ARC-2? It's because ARC-1 played its role quite well. And in particular, it was able to capture this shift towards AI systems that, for the first time, had actual fluid intelligence, which was completely new.
François Chollet [5:27] GPT-5 does not have fluid intelligence, for instance, but o3 does. That's a huge difference. And it was the only benchmark that caught that. But the benchmark still had a number of flaws. The ideas behind it were solid, like the idea, for instance, that every task would be completely unique so that you cannot prepare for it in advance; the idea that all the tasks would be built only on top of core knowledge priors, meaning that they only require the kind of knowledge that's extremely basic, that everybody has, that LLMs also have, of course, that any child has effectively.
François Chollet [6:06] But there were some issues with the execution. So in particular, not every task was 100% unique. There was not quite enough diversity in task space. And most tasks were actually very brute-forceable, meaning that if you would just look at the task, you could instantly see the answer without having to do much thinking. And that means that if you're using a brute-force program enumeration process, for instance, to try to find the program that will solve the task, you actually have a pretty good shot at finding it because the space is pretty small.
François Chollet [6:44] And so in ARC-2, we are fixing that completely. So now it's much, much less brute-forceable. Any task that you see is going to take you at least several seconds of deep thinking, maybe even minutes. So for instance, we actually ran all these tasks, thousands of them, by actual people in live, in-person sessions in San Diego. And so what we found is that everybody can solve these tasks. We only kept tasks that could actually be solved reliably by multiple people.
François Chollet [7:22] And the average time to solve one task was about five minutes. So it also tells you that they're not trivial. They're not super easy. It actually takes you quite a bit of thinking. So it's basically just an upgraded version of ARC-1. It's the same ideas, the same format, the same principles behind it, the same goal about what it's trying to measure: just machine fluid intelligence. But the execution is better.
Limitations of brute-force scaling in LLMs
Matt Turck [7:55] We are going to go into all of this and unpack all of this. But maybe taking a step back, you are one of the world's foremost experts on deep learning. However, historically, at least, you've been one of the few people, I would say, that in the context of massive hype around LLMs has had a much more nuanced view on LLMs, and in particular the brute-force aspect. So, to provide context to this whole conversation, what are LLMs good at and not good at?
Matt Turck [8:31] And maybe as a preview, we'll talk about this in a second. But LLMs up until o3, and we'll talk about this, have performed very badly on ARC-1. The early numbers seem to show that they are not performing particularly well either on ARC-2. So what are the fundamental limitations of the LLM approach, solely the brute-force approach?
François Chollet [9:02] Yeah, absolutely. And it's important to note that it's not like recently LLMs started doing well on ARC. LLMs are still doing extremely poorly on ARC. What's actually going on is that we've seen a shift from systems that were just base LLMs used statically at inference time to much more sophisticated systems that are no longer LLM systems that use LLMs as part of a search loop, but fundamentally they're not end-to-end deep learning models anymore. In general, LLM is a very overloaded term, and different people can mean different things by that.
François Chollet [9:42] There's of course the base LLMs, where you have a prompt and you're just doing one generation at inference time. Then you have the so-called reasoning models that are first going through this step of generating a chain of thought to improve on your prompt and try to adapt a bit better to the task at hand instead of just trying to one-shot the answer. And then you have models that actually perform test-time search, test-time adaptation. Some of them might be doing things like test-time training, test-time fine-tuning, which is not something you see in commercial models.
François Chollet [10:25] But it's a technique you can apply. And increasingly, you're going to see commercial models that use test-time search, where instead of just trying to generate one single CoT to adapt to the task, they're actually going to run through this search. It could be, for instance, MCTS-style search to try to synthesize the perfect chain of thought for the task. And so you really have to keep in mind that there are very different kinds of systems with different kinds of abilities. And the limitations of the base models are still the same.
François Chollet [11:01] And a benchmark like ARC, which was designed to highlight these limitations, has completely resisted the paradigm of scaling up these base models. Back in 2019, when I released ARC initially, the best LLM out there was GPT-2, which of course was completely useless on ARC. Then about half a year later, we had GPT-3, which was also doing 0% on ARC. And today the latest base LLMs, they're doing something like 10% on ARC-1, but on ARC-2, they're doing 0%. And to be clear, 10% on ARC-1 is very low because you would be doing like 95% plus, perhaps.
François Chollet [11:11] No, definitely. Maybe you should try the eval set.
Matt Turck [11:13] I did try a few.
François Chollet [11:40] It's not difficult, right? It's really straightforward, especially if you take your time, you take a few minutes per task. And even if you're just doing extremely crude brute-force program enumeration, like what we saw in the first competition in 2020, you get like 20%. If you scale up the same technique just a little bit, you get to like 50%. So if you just do 10%, that's extremely low. So since 2019, there's been this 50,000x scale-up of base models that has resulted in basically a flat curve on ARC.
François Chollet [12:13] So where we've made a lot of progress is with test-time adaptation. So your question initially was about what LLMs are good at, not good at. Well, if you're talking about the base models, they're kind of bad at everything, to be honest. But what you can use them for is things like non-critical information retrieval, for instance, or automating tasks where failure is not costly. Because the thing is that they may be able to do something useful for you, but there's always some probability that they will fail or they will hallucinate in some way.
François Chollet [12:49] So you should definitely be using them with a human in the loop and for tasks that are non-critical. And if you're using them for brainstorming, information retrieval, and so on, you should make sure that you check your outputs if you're going to use them for anything. But they're super useful for a very wide range of minor automation tasks. So I use LLMs daily for all kinds of tiny things, for coding, of course, as a sort of text-line completion mechanism, or a replacement for Stack Overflow as well, if there's a question I need to answer. For brainstorming, they're great as well, checking typos in a piece of text and so on.
François Chollet [13:25] Lots of things that they're really useful for. What they're bad at is everything that's more like algorithmic-style reasoning or adapting to a problem they've never seen before. And that's why they fail on ARC, for instance.
Defining intelligence: adaptation and efficiency
Matt Turck [14:11] As further background on all of this, so ARC is a benchmark that ultimately focuses on the progress towards AGI, which gets us into this whole conversation about what AGI is and what intelligence is. And that's something that you've been thinking, obviously, about for a number of years now. And so, related to those limitations of at least the brute-force aspect of LLMs, what is intelligence, and what kind of intelligence needs to be developed that this test-time compute thing is trying to get at?
François Chollet [14:50] Sure. I mean, again, intelligence is a very, very overloaded term, and different people mean very different things by it. But the definition of intelligence that I'm really interested in is not intelligence as the ability to do a bunch of specific things, basically a bag of skills. I'm interested in fluid intelligence, basically the ability to approach a new problem and creatively solve it, or approach a new domain and pick up a skill that you didn't have before, and do that efficiently. And I think the notion of efficiency is really at the heart of what intelligence is, because you can always do things expensively.
François Chollet [15:30] You can always use brute force, for instance, to solve a problem. Well, here's a very easy algorithm: just consider the space of all possible solutions and iteratively check them one by one. Eventually, in a billion years, you'll find the right solution. You can really apply that to any problem at all, including art, by the way. You can definitely solve art via brute forcing if you have billions and billions of dollars in compute. But that's not what intelligence is.
François Chollet [16:06] Intelligence is the ability to take a shortcut, to basically do these very hard things, solve these very hard NP-complete problems effectively, dramatically efficiently, like in linear time, effectively. And yeah, we are intelligent because we can learn fast, pick up new skills fast, solve problems fast. Like when you solve an ARC task, it takes you maybe a few seconds, maybe a few minutes. You don't have to iterate over 100 million different possible programs that go from the input to the output. You look at maybe two or three possibilities.
How O3 achieved a massive leap in ARC performance
François Chollet [16:23] And yeah, so to summarize that, I see intelligence as skill acquisition efficiency. So it's not the fact that you can acquire skills; it's how efficiently you can do it. That's a measure of your intelligence.
Matt Turck [16:47] And the ability to adapt, right? So how does that manifest then? So that's a lot of your work. And I guess o3 seems to be, like, a step in that direction. So, to put it in context, all the LLMs were pretty dramatically failing, as we were saying earlier, on ARC-AGI-1 for a bit. And I believe shortly after the end of ARC Prize 2024, you guys got a call from OpenAI, and OpenAI said, "Hey, actually, we're doing some work on o3, and it seems to be producing some interesting results."
Matt Turck [17:29] What is it, from what you know, that they did? And you can get into any level of technical detail you'd like, that got them from whatever, 20%? Claude 3.5 from Anthropic was the best at 25%, and then there was a jump. o3, I think, went to 75%, and 85% for the private?
François Chollet [18:04] On the small private set, which is what we actually used for the leaderboard, it was much, much lower. Generally, base LLMs were, like, at most around 10%. But yeah, so why this sudden huge jump? It's really because you have this paradigm shift away from just scaling up base models to doing test-time adaptation. And o3, best I can tell, is the most advanced, the most successful test-time adaptation model out there at this time. And what were they doing exactly? Well, I can only speculate.
François Chollet [18:37] Clearly it's doing test-time adaptation. That's why you're seeing this huge jump in generalization capabilities. But how does it work exactly? So my informed speculation would be that they're probably doing some form of test-time search over the space of possible chains of thought. And every time they try to make an edit to the current chain of thought, they're checking for self-consistency, for consistency with the task, and so on. And the check is performed automatically by something that you would call a grader model, a model that's just specialized in evaluating whether your chain of thought is actually solid, whether it's self-consistent, whether it's a good match for the problem, and so on.
François Chollet [19:28] So it's effectively a form of test-time program search, where the program is actually a natural-language program in token space, and where, instead of verifying correctness by actually executing it against, like, a suite of unit tests, for instance, you're actually asking another model to evaluate it for you. And yeah, that's what I think is the most likely, in the sense that it is consistent with what you observe, with the information that we have. And so, in general, one way you can tell apart a model that is just a base model, or maybe a model with a single chain-of-thought generation, versus a model that does this time search, is you can look at the inference latency.
François Chollet [20:13] If the model does search, it's going to take minutes to answer. Maybe it's even going to take longer. And you can look at cost as well. A model that does test-time search, it's doing a lot, lot more work at test time. It might cost 10x, 100x, 1,000x more. For instance, OpenAI o3 on the highest compute settings that we tried it on for ARC, it was consuming somewhere between $10,000 to $20,000 per task, for one little puzzle, which you could normally solve with a base LLM API for a few cents.
Speculation on O3's architecture and test-time search
François Chollet [20:38] And another thing you can look at is just generalization power. And ARC is actually a very good benchmark to probe generalization power. And so you see this huge jump from base LLMs to models that perform this adaptation, just massive. We're going from 10% to 80%.
Matt Turck [20:43] And do you know if they fine-tune on the dataset that was provided by—
François Chollet [21:08] That's a great question. And we don't have a clear answer there. So they told us that they were using a significant fraction, I think they said something like 75% of the training tasks, to adapt the model in some way. So of course, to start with, that's entirely legit. The training tasks are there to train on them. So if you're training on the training data, that's perfectly legit. That does not invalidate the score that they're getting on the semi-private test set in particular.
François Chollet [21:44] But what did it mean exactly? What were they doing with these tasks? There are at least two interpretations you can have. One is that, oh, they were just part of the pre-training data for the base model. That doesn't really make any sense because everything was part of the training data for the base model. In particular, that's a model that's definitely trained on GitHub. Well, our training data is definitely on GitHub, not 75% of it, but 100% of it. And plus the public eval data is also on GitHub.
François Chollet [22:14] So it would also be trained on that, definitely. And on top of that, it will also be trained on many, many different repos on GitHub that provide more generated ARC tasks or that provide solution programs to specific ARC tasks from the training set and the eval set and so on. So presumably that's not what they meant. They must have been doing something special with specifically 75% of the training tasks. So another interpretation is that they have a system that can do few-shot fine-tuning via reinforcement learning, for instance.
Program synthesis: what it is and why it matters
François Chollet [22:54] of their model. And so they created an instance of the model, of the base model, that's just specifically fine-tuned for ARC. I think that's what's most plausible here. And that would certainly explain parts of the performance jump. But still, these are capabilities that we've never seen in a model before. So whether that's a model that's just ARC-specific or it's a more general model, regardless, I think that's extremely impressive. It's definitely a breakthrough.
Matt Turck [23:33] And there were some other very interesting and promising approaches in ARC Prize 2024. Can you talk about those and, more generally, how that ties with your own work? So things like deep learning-guided program synthesis. What does that mean? What does program synthesis mean, and how does deep learning help? And then some of the other top approaches were test-time training and then combining program synthesis with transductive models. What does that all mean?
François Chollet [24:05] I think ARC Prize, last year, did a tremendous job at highlighting what are the current best approaches to create models that actually have fluid intelligence. And they're all test-time adaptation methods. So in particular, one category of method that really got big via ARC Prize is test-time fine-tuning or test-time training, where you're using an LLM that's doing transduction, meaning that it's looking at the task and it's trying to directly predict the answer. You can think of it as opposed to doing induction, which is that you look at the task and you try to predict a program that will turn the task into the answer.
François Chollet [24:44] So in one case, you just directly predict the answer. In the other case, you try to predict the process or program that gives you the answer. And so you're looking at these transductive models, and at test time, you're going to generate input-output pairs from the current task that you're trying to solve, specifically for fine-tuning your model to map one input to the output. And then you're going to run that model on the test input and see what it gives you.
François Chollet [25:17] So that's one thing. Another category of approaches that got big with ARC Prize is program synthesis. So program synthesis is this idea that you have some language that you're working with. Maybe it's a language that's specific to the problem, like a DSL, a domain-specific language. There are many ARC DSLs out there. Maybe you could also just use a generic programming language like Python. And at test time, you're going to look at the task and you're going to try to write a program that solves the task.
François Chollet [25:55] And then, in the process, you're going to generate a bunch of candidate programs, and you're going to try to run them on the inputs and see if they give you the right outputs. If you find a program that's correct, it actually works on the test pairs, and it's short, it's reasonably short, then it's a pretty good candidate to be a program that will generalize on the test input. And there are many different flavors of that idea. So you could use, for instance, an LLM to generate the code in Python.
François Chollet [26:23] That's something you would do. That would be one instance of deep learning-guided program synthesis. You can also do actual discrete program search, meaning that you have operations, like elements in the DSL, and you're trying to recombine them into graphs, which are programs, and you do this recombination just via a discrete search process. So that's actually the kind of technique that performed best in the prior editions of the competition. In fact, the very first edition in 2020 was won by a technique that was just doing this with a handcrafted DSL and a very, very cool search process.
Matt Turck [26:55] What are the particular technical challenges to program synthesis? So if you think of the LLM approach as being constrained by data, what is a constraint for program synthesis?
François Chollet [27:26] Well, the big constraint is that the way you write your program is not very sophisticated. So the probability that the first program that you try is correct is very, very low. To find the correct program, you're going to have to try many, many different programs. For instance, if you're just doing brute-force program search, you're trying programs at random. And so to find the right program, you have to iterate with millions of potential solutions or points in the search space. And the main bottleneck is that this search process takes a very, very long time.
François Chollet [27:59] It's a very large search space. And to evaluate all these points, each point takes you some amount of computation to evaluate. You're going to have to evaluate millions of points. This is extremely expensive. And of course, it's not how humans do it. Humans probably do something akin to program search, but they're doing very, very little search. You look at a problem, you're quickly building a very small set of hypotheses to explain what you're seeing. Maybe you have two, maybe you have three, and you test them.
Combining LLMs with search and synthesis techniques
François Chollet [28:34] So it's a form of search. It's a form of program search with backtracking. If one of your programs doesn't work, you can just mentally debug it or just discard it and move on to the next hypothesis. But the size of the search space is tiny because humans are fundamentally not very good at these things that computers are extremely good at, like evaluating discrete search spaces with millions of points, for instance. So humans can have some shortcuts to find the right program intuitively.
Matt Turck [28:56] That sounds like the fundamental concept of intuition. So is part of the idea of what you're doing and what you're suggesting to combine brute force again, or deep learning for what it's good at, and then try to bring that concept of intuition as our path to AGI?
François Chollet [29:20] Yeah, I mean, we don't know how humans do it, but clearly intuition is part of the picture. So intuition is this idea that you're going to leverage the experience you have to make fewer guesses, so that there are fewer guesses that you have to check before you find the correct guess. A very crude version of that is what LLMs are doing. With an LLM, when you generate a program conditioned on some task definition, effectively you're using a statistical prior about the shape of program space to generate fewer guesses, right?
François Chollet [29:46] So most of your guesses are going to be syntactically correct, for instance, which is already a nice property. Most of them are going to be related to the task in some way, and so on. And so if you're using a statistical prior like this, maybe you can find the right program with a few million programs, or even less than that, a few tens of thousands, maybe thousands of programs, instead of trying billions and billions, which is what you would have to do if you were just doing brute-force program enumeration.
François Chollet [30:31] But the more sophisticated your intuition is, the fewer guesses you have to make to find the right answer. And if you look at an extremely sophisticated intuition system like humans, we just see the right answer. We only make very, very few hypotheses in our mind, like just a handful, effectively.
Matt Turck [30:51] Just to make sure I understand, is program search and program synthesis, are those the same things, or are they different concepts? In other words, are you just trying to select in a long list of existing programs, or are we talking about basically AI creating a program to adapt to a situation it hasn't seen before on the fly?
François Chollet [31:12] They're usually the same thing because if you're doing any kind of program synthesis, even by using an LLM to generate the program, for instance, you still have to check that it is correct. It's probably not going to be correct on the first try. So if you're doing program synthesis, you're definitely doing some form of program search. You have to test a few guesses before you find the right one.
Matt Turck [31:33] All right, François, thank you so much for this. A wonderful time to bring in your co-founder, Mike, who was previously the co-founder of Zapier. First of all, how did you guys connect? How did you guys start working together? It was a mutual friend who introduced us, Lukas Biewald, the Weights & Biases CEO. My background, like you mentioned, co-founded Zapier. I've been basically running that company for the last over 10 years and started getting into AI more deeply in the beginning of 2022.
Matt Turck [32:01] The chain-of-thought paper that came out in January that year was one that really shocked me. I'd been paying attention, loosely curious about AI, basically since the beginning of the company. But Zapier is an automation company, not really necessarily investing in the frontier of deep learning and machine learning. I gave a whole presentation at the company on GPT-3 and saw this chain-of-thought paper that came out over a year later and was really shocked that all these reasoning benchmarks that we had at the time, these language model benchmarks, were showing these step-function increases in score using this step-by-step, "let's think step by step" method.
Matt Turck [32:44] I was running half the company at the time. I gave it all back to my co-founder, Wade, who was the CEO, and basically just went all in on doing AI research at Zapier. I think that was one of the things that led Zapier to being a really early adopter of AI and deployer of AI into the market. We've been deploying AI agents now for over two years. One of the interesting things that I noticed for years was the feedback from customers.
Matt Turck [33:07] I've spent time with hundreds of thousands of them talking about trying to use AI to do automation for their businesses. The feedback was always the same, which was like, hey, I get the promise of this stuff. I know what I want to try and use it for, but it's not working reliably enough. It fails randomly two out of 10 times, which might be fine in a supervised setting like ChatGPT, where you're talking with these assistants, but it doesn't really work in an automation case where it's hands-off keyboard, running on a server, where you're not actively supervising it.
Matt Turck [33:50] This was sort of the same consistent feedback even as the models were getting better. I kept hearing the same feedback. This was at the same time of all of this AGI scaling hype that I experienced, and I'm sure you did as well, during 2023 and 2024 online on Twitter. So I had these two sets of lived experiences that didn't mesh well and was trying to start figuring out what the truth was, and came across François's paper. I think I first got introduced, François, to your work back during COVID on the Lex Fridman podcast and rediscovered it during this period at Zapier.
Matt Turck [34:23] I started just surveying people in the AI landscape around the Bay Area, like, hey, have you heard of this ARC benchmark? It seems really important. I think it might actually be the most important unbeaten benchmark that exists today. And barely anyone had heard about it. So finally, I was chatting with Lukas. It turned out he knew François really well and introduced us. I flew up to Seattle actually just literally one year ago and pitched him on some ideas of how to try to beat the benchmark, as well as check my understanding.
The ARC Prize structure: efficiency track, private vs. public
Matt Turck [35:01] Like, hey, is it truly low awareness? Do you also agree that it's a really, really important unbeaten benchmark? I think I walked away from that conversation believing the answer was yes to both those questions. I had just seen Nat Friedman run this Vesuvius Challenge the year prior, and it was really, really successful. It was one of these competitions to focus a lot of attention on a problem that most people just never heard about or were not aware of. I felt like we could emulate that, and ultimately that was kind of how we came to launch the initial version of ARC Prize together.
Matt Turck [35:36] And let's get into the specifics of how that actually works. So, starting with 2024, but I guess the structure is the same today. So there's a private version, there's a public version, there is a paper prize. How does that all work? I think some of the goals that we have for the ARC Prize are to—one initially was to raise awareness of the fact that there was an unbeaten, really important AI benchmark out there that wasn't saturating. It resisted this 50,000x scale-up in pure language model systems.
Matt Turck [35:55] This was to counter and provide some public education of the fact that pre-training alone is not enough for these pure language models. This was in contrast to, again, all of the hype, all of the dogma that I think you saw online. Just to put a really clear point on this, if you don't really remember this era, maybe remember last summer in California, there was that big SB 1047 bill, that AI bill that was primarily regulated and introduced on the thesis that this scaling was going to continue and lead to really, really bad situations, and we have to impose regulation right now because if we don't right now, it's going to be too late.
François Chollet [36:46] Yeah, it's like back when GPT-4 was released, the universal narrative, like everywhere in the tech industry, was, okay, GPT-4 is just a bigger version of GPT-3 trained on more data. So it's the same architecture, it's the same principles behind it. It's just scaled up, and look at this incredible step-function improvement. And GPT-3 itself was just a bigger version of GPT-2. And so the idea was that, well, GPT-5 is just going to be the same, but it's going to be 100x bigger.
François Chollet [37:20] And we're going to need these massive data centers to train it. And it's going to be truly AGI. AGI is just going to spontaneously emerge from scaling up this thing. And basically all the evidence was pointing in this direction. If you looked at benchmarks that were based on just brute-force memorization of knowledge and skills, if you didn't look at ARC, this was the takeaway you had. And so everybody at the time believed that. And maybe some more sophisticated researchers, like maybe some at OpenAI, I don't know, maybe some at DeepMind, had different beliefs, but that was really the mainstream narrative.
François Chollet [37:53] And, well, interestingly, one year later, like about one year and a half later, this narrative was gone. People were completely aware that no, just scaling up pre-training is not all you need, that you actually need entirely different ideas and maybe, in particular, test-time adaptation to achieve actual fluid intelligence and AGI.
Matt Turck [38:16] So the first big goal is to raise this awareness of this fact. And honestly, it was to re-inspire folks. I think I'd spent a lot of time with researchers and students over the year prior to that. And one of the interesting sentiments I heard was just a lot of dejection. Like, hey, isn't all this stuff figured out? There's really not that many interesting things to do in AI. Maybe I should go work at the application layer on language models instead of the research layer.
Matt Turck [38:45] I felt like this was a really not-good thing. We don't have AGI yet. We don't have the ideas for it yet. We still are idea-constrained, even literally today. If that's true, which I think ARC-AGI-1 and ARC-AGI-2 both show it is, you want to design the strongest innovation ecosystem you possibly can, which is going to be one that's very open. There's a lot of sharing, there's a lot of diversity of approach. What you don't want is a very closed-up environment with no sharing, where there's dogma or a monocultural viewpoint of how to do it.
Matt Turck [39:10] And so we really wanted to try and inspire folks with our prize last year to work on new ideas. And to this point, one of the versions of the private version, correct me if I'm wrong, you have to publish exactly what you do. There's a limit on the amount of compute. Is that fair?
François Chollet [39:38] Yeah, I mean, the basic idea is that there's one track for self-contained approaches with no internet access that are very efficient. So they have a very limited compute budget, and authors must open-source them at the end of the competition. So this is really designed to incentivize open sharing and get as many ideas as possible, with a big focus on efficiency. We believe efficiency is not just a good feature to have in your system. It's actually at the heart of intelligence.
François Chollet [40:10] And the other track is to provide continuous benchmarking of frontier models to be able to track, okay, if you look at the best currently available commercial frontier models, things like right now, for instance, o1 Pro, soon in the future it's just going to be o3 and so on, Gemini 3, whatever. How much fluid intelligence do these models actually have? Mostly independently from the efficiency consideration. We still do have this notion that we want to monitor efficiency. So we're going to be reporting results on a 2D plot.
François Chollet [40:30] So we're not just looking at the score as a scalar, we're looking at the score associated with the cost per task that was required to achieve this score. And of course, a model that gets you the same score, but at much lower cost per task, is a smarter model.
Matt Turck [40:57] One other thing we've got too, Matt, on the contest that's worth adding, just on the structure of it, is we've got these score tracks that François was just talking about. We also introduced something last year, which we're doing again for 2025, which is this paper award track. One of the critiques of things like Kaggle contests is you get a lot of overfitting to the dataset or to the benchmark. People do dataset mining, lots of very narrow approaches that maybe don't generalize because they're optimizing just to win the contest, win the top score.
Matt Turck [41:26] Certainly, we actually saw some of that. There were some legitimate good conceptual breakthroughs as well, but certainly there's mixed in a lot of just hacking the benchmark. We had to introduce this paper award alongside it in order to basically incentivize conceptual progress. I actually think some of the best research that came out from ARC Prize 2024 was actually on the paper track last year. Things like these test-time training, test-time adaptation approaches, we actually got full, not only open-source reproducible code, but the theory that backed it too.
Matt Turck [41:59] I think one thing that potentially ARC Prize 2024 will be remembered for in retrospect is going to be marking a moment in time where AI realized, ah, yes, we do need test-time adaptation methods to be able to solve things like ARC. And that's certainly what o3 shows as well. We spoke with François a little bit about the o3 aspect. But separately, the top result was—was it MindsAI? There was Jack Cole, that group.
Open source as a requirement for progress
François Chollet [42:10] So, yeah, they were at the top of the leaderboard up until the end of the competition. They ended up dropping out because they did not want to open-source their solution. And of course, that meant that they were not eligible for the prize.
Matt Turck [42:31] On that topic, open source is a big aspect of what you do. Maybe talk to this and how this year, 2025, is different from last year? Last year, there was absolutely a narrative that OpenAI was so closed-source that they were slowing things down in the world, in the progress towards AGI. This year, we're in a DeepSeek kind of moment. What do you make of all of this? I've had a lot of teams reach out and tell me they're very excited to actually use and try to use some of the new DeepSeek-like systems and the ideas from the paper actually on ARC Prize.
Matt Turck [42:56] One of the goals with open source, really, two of them. Why do we require open source for the prize? There's really two reasons. One, we were trying to emulate a bit more of what the academic environment around AI research looked like during the 2010s to 2020s. If you look at even where we got to with language models, that was built on open progress, open sharing between four or five different people across four or five different companies and labs that ultimately led to the Transformer, into GPT-2, and the scale-up from there. So we wanted to try and emulate a bit more of that.
Matt Turck [43:36] Unfortunately, just due to a lot of the competitive market dynamics, basically since GPT-4 came out, there really was no open idea sharing at the frontier. That felt like it was really setting progress back in the field to us. So we wanted to try and use this prize as a way to, in a small way, shift the frontier back a little bit more towards openness and get a little bit more sharing done. And then the second thing we added is in order to be able to have the community consistently get rebaselined at the end of the contest going into next year.
Matt Turck [44:10] One of the features and aspects of contests is, while we hope that people would share, the reality is people want to hold their alpha. They want to not share during the contest in order to win the prize at the end. One of the structural designs of the contest is we've been running them annually. We'll continue doing it annually until the grand prize is beaten. During the contest, people aren't really going to share. We know that. But then we hopefully can use prizes at the end of the contest in order to convince people to share progress, use that open knowledge then to rebaseline the community around what it takes to actually beat this thing, and use that knowledge going into the next year then.
Matt Turck [44:47] In a funny way, it's kind of ended up being what happened this time. We got a bunch of test-time adaptation methods that came out of ARC Prize 2024. Then we got this big set of evidence from o3 that test-time adaptation is a legitimate way that you can make significant progress on the benchmark, but maybe not very efficiently. Then we had the DeepSeek stuff come out in January. All of this has now entered the frontier knowledge set of AI researchers going into ARC Prize 2025.
What's new in ARC-AGI 2 and human benchmark testing
Matt Turck [45:17] We're really excited to see people bring and apply some of those ideas to the new ARC-AGI-2 dataset this year. And again, for anyone that hasn't been exposed to it, you have to get to 85%. 85% is sort of the success threshold where you actually win the prize. ARC Prize 2025, there's a bit of a changelog from last year. The big one is we're swapping version 1 for version 2 of the datasets. We're using ARC-AGI-2 now. You have to get 85% within Kaggle's efficiency limits. And 85%—is it human-level solving those problems?
Matt Turck [45:47] It's actually lower. One of the cool things about the new ARC-AGI-2 dataset is we did a big controlled human study down in San Diego over the winter. We can strongly assert that every single task in the dataset was solved by at least 2 humans in under 2 attempts, which is the same rules that we apply for the AI systems as well. This was actually not something we were able to say confidently about version 1. We believed it was true anecdotally, but it wasn't something we had data for.
Matt Turck [46:17] Now we actually have data to say this. We said 85% mostly because we expect that there's some degree of ambiguity in some of the tasks. There's probably some degree of small bugs in the tasks, and we just didn't want to hold the 100% bar as the target. We felt like the 85% bar was a good one that recognized that there's probably some fuzziness around the edge here. But every single task in this dataset has been solved by humans.
François Chollet [46:42] To be clear, an average person on their own would not score 85%, probably. Based on our own testing, an average person in our test sample would score about 60%. But if you take a small panel of people from our test sample, like about 10, and you just have them solve the task independently and then do majority voting, that panel should score collectively 100%.
Matt Turck [47:08] Going into what is new with ARC-AGI-2, maybe let's talk about the panel. Obviously, the question—and we talked about intelligence and AGI a little earlier—but there is a spectrum of intelligence across humans. How did you put the task together in a way that would reflect that spectrum of intelligence? How do you select people? What was the process?
François Chollet [47:36] We were not trying to select for people that were good at solving puzzles, for instance, that would be strong at ARC. We really hired random people. We went through a service that exists to basically just hire people for scientific testing. And we ended up with a pretty diverse crowd. We have anywhere from Uber drivers, students, people who are unemployed. So basically anyone trying to make some money on the side. So very much regular folks. And we know our tasks are solvable by regular folks.
François Chollet [48:00] One thing that we did with the task selection process is that we tried to eliminate tasks that could be solved by everyone, because if something is just universally feasible, it doesn't actually give you a very strong signal. It's probably something that can be brute-forced very easily. And we also eliminated tasks that were too hard. So we only kept tasks that could be solved by at least 2 people independently in under 2 attempts so that we know that it's solvable and it seems, in fact, reproducibly solvable.
François Chollet [48:13] It's more than just a one-off.
Capabilities ARC-AGI 2 is designed to test
Matt Turck [48:31] I noted in your write-up that there were certain capabilities that you were focusing on, including symbolic interpretation, compositional reasoning, and contextual rule application. What do those mean?
François Chollet [48:39] We can go into what each one of them actually represents, if you'd like.
Matt Turck [48:41] Or directionally.
François Chollet [49:13] The high-level takeaway is that the new ARC tasks all require some level of deliberate, deep thinking. You have to look at the data, you have to think for a while, and the reasoning chain that you're going to have to come up with is going to have a few hops. In particular, it might have a few control-flow hops, like something like an if statement, for instance. And that's something that we're not seeing a lot in ARC-AGI-1. And it really creates this big difference in terms of how challenging the dataset is for LLMs in particular.
When will ARC-AGI 2 be saturated? AGI timelines
Matt Turck [49:35] When do you expect ARC-AGI 2 to be saturated? Is the idea that there's going to be a 3 and then a 4? And at what point can we collectively say a big hurrah, we've reached AGI?
François Chollet [49:53] Right. I mean, pragmatically, you have AGI when it's no longer possible to easily come up with tasks that you and I can do naturally, but no AI system can do. If it's no longer possible to come up with this kind of task, then you probably have AGI.
Matt Turck [50:02] I was surprised, by the way, how easy it was for us to come up with the V2 dataset. I think that actually is somewhat informative about how far we have to go still.
François Chollet [50:24] Yeah. The closer you are to AGI, the harder it's going to be to come up with this kind of problem. And right now, it's not difficult to come up with this kind of problem just yet. So we still have some time to run. So how long will it take for ARC-AGI 2 to be saturated? We don't know. And it's not just a question of when will, like, 85% be reached? It's also the question of how efficiently it's going to be achieved.
François Chollet [50:50] But we're already starting to work on version 3, which should have a brand-new format. So we don't think ARC-AGI 2 is going to be the last benchmark. I think that there's definitely going to be a next benchmark after that, and then a next benchmark after that. There's probably still a couple of important milestones on the way to AGI.
Matt Turck [51:15] I guess ultimately, unpredictable, because surprises come up. I think O3 was a surprise, at least for me personally. I think this is one of the things that a lot of forecasters in AI and commentators and even people reading it struggle to recognize, which is you can make pretty easy predictions along smooth scaling curves. You just draw the trend line and say, okay, here's where we expect to get to. The much harder thing to predict is when step-function changes occur, because these are not the results of a smooth scaling law that you're applying something into or tweaking a variable along.
Matt Turck [51:56] This is something where it's like, hey, the underlying dynamics and structure of how these systems work has actually changed. It requires some sort of new idea that got forged in. One of the really tough things about making predictions like that is two things are necessary. One, the idea must be in the world. And two, someone has to implement the idea. And in an environment where that hasn't happened yet, you're actually not sure which of those two things is true. Is it like, oh, we just don't have the ideas yet, and the technology hasn't caught up enough for us to even produce the right idea of how to build it yet?
Founding of NDEA and why now
Matt Turck [52:32] Or is it like, no, all the knowledge is kind of out there, and someone needs to put it all together into an implementation? So both of these are very inherently unpredictable things that are hard to ascertain until, in retrospect. It makes forecasting around this stuff quite challenging. And I'm always somewhat reticent to put hard stakes in the ground of when will step-function changes happen. Those tend to be very surprising moments, and I think '23 was certainly one of them.
Matt Turck [53:06] I'd love to talk a little bit about Ndea because you guys are co-founders of the ARC Prize, but you're more recently, and perhaps more importantly, co-founders of a new research lab. So I know you cannot talk about it too much, but whatever you can share, including the idea itself. François, you've had an incredible ride at Google. You're famously the creator of Keras. And there are a lot of people that decide to stay in the comfort of those large companies, and some people start a company and decide to go back to those companies.
Matt Turck [53:17] What was the reasoning behind partnering with Mike and venturing out on your own to pursue AGI?
François Chollet [53:44] Well, we have a vision about what we want to build, and it's a vision that's very different from what the other frontier labs are actually looking into. So we have some insights, we have some beliefs that are going against the mainstream narrative in the AI research community. And we believe the best way to actually execute on that is as an independent small-scale entity. What's also going to be very important for us is to move fast, to iterate fast, and that's much easier to achieve as a small, scrappy team.
Matt Turck [54:02] At least a direction, based on the little bit that you guys have written on your website, is around precisely that concept of program synthesis?
François Chollet [54:17] That's right. We really believe that the frontier is test-time adaptation and that the best approach to solve test-time adaptation is deep learning-guided, or intuition-guided, program search. And so that's what we are trying to achieve.
Vision beyond AGI: a factory for scientific advancement
Matt Turck [54:47] You mentioned that actually getting to AGI was a part of it, but actually the vision was even bigger than AGI. And to quote, you wrote that you were building a factory for rapid scientific advancement, a factory capable of inventing and commercializing NDEs, hence the name. Any additional color to help contextualize what that means? I think it was the thing that we found a lot of sort of values alignment around last summer. We were working on ARC Prize together. Both of us, I think, got really excited about trying to apply the frontier of AGI towards trying to compress scientific timelines.
Matt Turck [55:25] I think there's kind of an interesting observation that basically, if you sort of zoom out and ask, what is humanity today? Humanity is not sort of like the collection of biological evolution that humans have gone over for the last 10,000 generations. It's really this knowledge colossus, this technology colossus that we've been able to build collectively and add on top of. AGI is going to dramatically have a chance to accelerate that rate of adding to that tech tree, that knowledge colossus. This is one of the reasons that's motivating even work on AGI, by the way, in the first place.
Matt Turck [55:58] If all we have and all we really get to in the near term, let's call it five, 10, several decades, is AI that is able to mimic and follow what humans have done, we're not really dramatically changing the rate at which we can produce new technology, produce new knowledge, to add it into that pile and give it back to humanity. You're always going to be constrained by the human that's in the loop that's reviewing it and guiding it and checking it. If we actually want to significantly increase the rate of change on this stuff, you really need AGI that can autonomously innovate.
Matt Turck [56:26] That's what's missing from today. That's really what's missing from the frontier. I think that's even in a small way what things like ARC-AGI-2 show is missing. If we can actually make progress towards this, we're going to make progress towards AGI systems that are capable of that autonomous innovation as well. So I think in our long term, we'd love to see Ndea become basically the most innovative company in the history of the world, producing the most number of new technologies, producing the most number of new knowledge using the technology we directly built.
What NDEA is building and why it's different from LLM labs
Matt Turck [56:55] So yeah, in a funny way, creating AGI is step one along a several-step master plan. And it's a scientific discovery, right? So it doesn't sound like your vision of success is helping people write better emails or do memes.
François Chollet [57:26] What is interesting to us personally, but it's also the fact that the kind of technology we are building, it's differentially advantaged, is that it's going to be capable of solving problems that have never been solved by humans before. That's a very different deal than LLMs, for instance. This makes sense, but effectively only in verifiable domains, like symbolic domains, things like mathematics, physics, programming, but not, for instance, creative writing. So we're not trying to help people write better emails. We think LLMs are actually a great system to achieve just that.
François Chollet [57:32] That's not what we're building.
Matt Turck [57:55] I kind of get excited about, all right, I'll add this. Look, AGI is going to get used to solve a lot of problems. We're already using AI to solve a lot of problems. In my experience, just seeing it at Zapier, Zapier's making money from these systems. So I have nothing bad to say about the current state of AI in terms of helping people solve important problems. We're going to use AGI to solve problems too. And I think that's totally fine, and we should absolutely do that.
Matt Turck [58:22] I do think that there's something really exciting about the creation, maybe discovery, of AGI in the sense that we're sort of going on an adventure here. There's a bit more of a journey into this unknown, unknown future. We're not exactly sure what the knowledge is going to be or what the technology is going to be yet. I think that's one of the things that personally gets me quite excited. It's almost like a bit of creating a time machine into the future in some way.
Hiring and remote-first culture at NDEA
Matt Turck [58:52] I do think that some portion of humanity's collective effort should be going and trying to push this adventure forward. It sounds like you have your core research team in place. Are you still presumably hiring? Basically, the founding team is in place. We're trying to basically build the top industry program synthesis team in the world. We think that's what it takes to be successful on the research program that we've got together. So we've got the founding team in place.
Matt Turck [59:17] However, we will always and forever consider exceptional candidates who are really great at program synthesis. So if that's somebody listening to this, you can apply. You can go to NDEA; there's a join link on the website. You can send us an email. We'd love to read it. And it could be anywhere in the world? Yes. Good. Thank you. You're doing a better job of promoting us than I am. Yes, NDEA is a globally remote team.
Matt Turck [59:38] This is kind of due to the distributed nature of where program synthesis talent lives. It's somewhat of a pragmatic thing. I don't know, there's maybe a couple million deep learning engineers now in the world. In contrast, there's probably maybe only a few hundred really great program synthesis folks across the world. And so, in order to be able to reach those folks and bring them into one team together, yeah, we're making this sort of pragmatic decision here to run a globally remote team.
Closing thoughts and the future of AI research
Matt Turck [1:00:14] And that's basically what Zapier—we're kind of following the footsteps of what I've been running for the last 15 years with Zapier, which is also globally remote. Wonderful. All right, well, that sounds incredibly exciting. I am, like a lot of people in the industry, very curious to see how you guys go with Ndea. Congratulations on everything you've achieved with the ARC Prize, and thank you for all the work that you're doing in the industry. It's been fantastic to see all of this.
Matt Turck [1:00:23] And most importantly, thank you for spending time with us today. I really appreciate it. Thank you. Thanks, Matt.
François Chollet [1:00:24] Thank you, Matt.