
The conversation
In a paper submitted 5 November 2019, On the Measure of Intelligence, researcher François Chollet proposed defining intelligence as 'skill-acquisition efficiency' rather than performance on any specific task, and introduced the Abstraction and Reasoning Corpus, ARC, grid-based puzzles meant to test that efficiency directly. In 2024, Chollet and Mike Knoop launched the ARC Prize, a competition built on a successor benchmark, ARC-AGI, drawn from Chollet's corpus. The prize's own site, arcprize.org, and its history page, describe the project's evolution as retrieved on 16 September 2026.
What the documents show
Chollet's paper argues that measuring performance on a fixed task, such as a board game, rewards a system for absorbing task-specific prior knowledge rather than for genuinely acquiring a new skill efficiently. ARC was built with priors closer to human cognitive foundations, so a puzzle can be easy for humans and hard for AI without specialised training. The prize's history page traces what followed: a first Kaggle competition in 2020, whose winner reached a 20% success rate; ARCathon competitions in 2022 and 2023, reaching roughly 30% by 2023; and the 2024 ARC Prize itself, where the year ended with a 'top score of 53% on the private evaluation set' while the grand prize remained unclaimed. The site also states that by 2026, 'all four frontier labs' report ARC-AGI scores on their model cards, and that the project introduced ARC-AGI-3, now testing agentic behaviour in interactive environments rather than static puzzles.
The system boundary
Chollet's own paper frames ARC as one proposed, narrower measure, contrasted with the broader idea of general intelligence it is often used to stand in for; scoring well on ARC-AGI demonstrates sample-efficient skill acquisition on this style of puzzle, not general reasoning across arbitrary domains. The prize's own leaderboard separates results by verification status and cost per task, distinguishing a verified run from a self-reported one, and cheap solutions from expensive ones solving more puzzles at far higher compute cost.
Where it fails
The 2024 competition's outcome undercuts a simple 'solved' narrative: the year's best system reached 53% while the grand prize, requiring a much higher bar, went unclaimed, and Chollet and Knoop's own framing states ARC 'shows us we still need new ideas' rather than that the format has been beaten. Because the benchmark evolves across versions, ARC-AGI-1, ARC-AGI-2, and now interactive ARC-AGI-3, a score on one version is not evidence about a score on another. A builder should treat a headline number as tied to a specific version and cost budget, not a fixed measure of general capability.
- Which ARC-AGI version does a cited score refer to, static puzzles or the newer interactive format?
- Was the reported result verified under the prize's own testing policy, or self-reported by the model's developer?
- What did the run cost per task, and does that cost make the approach practical outside a competition?
The benchmark's own history is one of a hypothesis repeatedly tested and only partly confirmed: efficient skill acquisition has proven harder to reach than raw task performance, the distinction Chollet's original paper set out to measure in the first place.
Sources & reading trail
Chollet's own proposed definition of intelligence as skill-acquisition efficiency and his introduction of the ARC dataset.
Source published: 5 November 2019 · Retrieved: 16 September 2026
The foundation's own timeline of ARC-AGI competitions, the 2024 top score and unclaimed grand prize, and the ARC-AGI-3 launch.
Source published: Not established · Retrieved: 16 September 2026
Documentation, rulings and incident records establish the entry; the boundary reading is Chatbot Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.