Datasets & Benchmarks · 2019
ARC-AGI
Francois Chollet designed puzzles a child can solve, then watched five years of AI progress barely dent them.
In November 2019, Francois Chollet, creator of the Keras deep learning library, published a provocative paper called On the Measure of Intelligence. Its argument: skill at any single task can be bought with enough data and compute, so benchmarks of skill measure the budget, not the intelligence. Real intelligence, he proposed, is the efficiency of acquiring new skills.
His benchmark, the Abstraction and Reasoning Corpus, embodies the argument. Each task shows a handful of colored grid pairs, an input and its transformed output, and the solver must infer the hidden rule and apply it to a fresh input. Every task uses a novel rule, so nothing can be memorized in advance.
The tasks assume only the core knowledge humans have by early childhood, notions like objects, counting, symmetry and simple geometry. Most people, including children, solve typical tasks easily. Machines did not: a 2020 Kaggle competition topped out around 20 percent using brute-force program search, and for years the best large language models scored close to zero even as they mastered exams like MMLU.
That stubbornness became the benchmark's message. Through the scaling boom, ARC stood as public evidence that bigger models were not automatically becoming general reasoners. In June 2024 Chollet and Mike Knoop launched the ARC Prize, offering over a million dollars to push progress on what they renamed ARC-AGI.
Movement finally came in December 2024, when OpenAI's o3 reasoning model scored 87.5 percent on the semi-private evaluation in a high-compute mode reportedly costing thousands of dollars of computation per task. Chollet called it a genuine breakthrough while noting the efficiency gap with humans, and released the harder ARC-AGI-2, keeping the benchmark's role intact: a moving target that measures generalization rather than accumulated skill.
From history to production
We turn these ideas into working systems
The same techniques, shipped into your stack with evals, observability, and measurable ROI.