Timeline

ARC Prize announces ARC-AGI-2 and ARC Prize 2025

The new 1,000-task benchmark reported single-digit scores for public reasoning systems, versus OpenAI o3's 75.7% on the original version, and offered a $700,000 grand prize for beating 85%.

  • Benchmarks & progress
  • Notable

The ARC Prize Foundation released ARC-AGI-2, a second-generation version of its Abstraction and Reasoning Corpus benchmark, alongside ARC Prize 2025, a $1 million competition hosted on Kaggle running from late March to early November. The original ARC-AGI, designed by François Chollet as a test of fluid reasoning that resisted memorisation, had come under strain months earlier when OpenAI’s o3 system scored 75.7% on it — a jump the foundation said showed the benchmark could be beaten by compute-intensive search rather than the general reasoning it was meant to isolate.

ARC-AGI-2 comprised 1,000 training tasks and 360 evaluation tasks, deliberately redesigned to remove puzzles vulnerable to brute-force approaches and to target weaknesses — symbolic interpretation, compositional reasoning, contextual rule application — that the foundation judged current systems handled poorly. Every task was calibrated against more than 400 human testers to ensure it remained solvable by people while pure language models scored zero and the strongest public reasoning systems scored only in the single digits. The benchmark also introduced cost-per-task reporting alongside accuracy, treating computational efficiency as part of what “intelligence” meant to measure rather than a separate concern.

The prize pool split into $125,000 in incremental progress awards, up to $175,000 in further prizes, and a $700,000 grand prize contingent on a system exceeding 85% while operating within an open-source, efficiency-constrained format. The relaunch reflected a recurring pattern in the field: a benchmark designed to be hard for AI systems is beaten faster than expected, and its successor is built harder still. The foundation published a fuller technical report on the benchmark’s design two months later.