ARC Prize publishes ARC-AGI-2 technical report
Humans solved all 1,417 test tasks in a median of under three minutes each; no frontier reasoning model exceeded 5% at launch.
- Benchmarks & progress
- Minor
The ARC Prize Foundation published the technical report for ARC-AGI-2, a successor to François Chollet’s 2019 Abstraction and Reasoning Corpus designed to remain hard for language models even as they improved on the original. Like its predecessor, the benchmark presents grid-based visual puzzles that require inferring a transformation rule from a handful of examples and applying it to a new grid — tasks meant to need little prior knowledge but genuine, on-the-spot abstraction rather than pattern retrieval from training data.
The report described three design goals: minimising prior knowledge required, resisting brute-force search, and calibrating difficulty against measured human performance. Tested on 400 people across 1,417 tasks, every task was solved by at least two participants, in a few minutes on average. Frontier reasoning models available at publication solved fewer than 5% of tasks — the report did not give individual scores — a wide gap from ARC-AGI-1, where some systems had already scored highly by spending heavily on inference-time compute.
The report mattered less for its numbers than for what it said about benchmark design: that a test built to isolate fluid intelligence from memorised knowledge could still open up a large, durable gap between human and machine performance, at a moment when many other benchmarks were approaching saturation. It also set the terms — and the near-zero scores — that subsequent frontier releases through 2025 and 2026 were measured against as they closed the gap.