Timeline

o3 posts a breakthrough score on ARC-AGI

A low-compute configuration scored 75.7%, roughly matching the ARC Prize's human-performance threshold, at about $26 per task against roughly $5 for a human solver.

  • Benchmarks & progress
  • Models & capabilities
  • Major

The ARC Prize Foundation, which runs the ARC-AGI benchmark designed by François Chollet to resist memorisation by testing novel visual-reasoning puzzles rather than recalled facts, published its independent analysis of OpenAI’s o3 model, tested ahead of public release. On the benchmark’s semi-private evaluation set, a lower-compute configuration of o3 scored 75.7%, and a much higher-compute configuration scored 87.5%. On the public evaluation set, the two configurations scored 82.8% and 91.5% respectively.

The scores represented a sharp jump from prior results: the foundation noted that ARC-AGI’s first version had taken roughly four years to progress from 0% with GPT-3 in 2020 to about 5% with GPT-4o in 2024, a trajectory o3 overtook in a single release, at a level close to the benchmark’s rough proxy for typical human performance. The cost of reaching that score varied enormously between configurations. The high-efficiency run, scoring 75.7% on the semi-private set, cost about $26 per task on average; the low-efficiency run, achieving the higher 87.5% score, used roughly 172 times as much compute, at around $4,560 per task, against an estimated $5 per task for a human solver — a gap ARC Prize highlighted as evidence that o3’s gains came partly from spending far more test-time compute per problem, not only from a more capable underlying model.

Chollet, while calling the result significant, was explicit that it did not mean o3 had reached general intelligence: he wrote that the model “fails on very easy tasks,” indicating its failure modes remained different in kind from a human’s, and the foundation said it would release a harder successor benchmark, ARC-AGI-2, designed to again separate genuine reasoning from pattern-matching at scale.

The result, announced the same day as OpenAI’s own o3 preview, intensified a debate already underway following o1’s earlier, much lower ARC-AGI score: whether test-time compute scaling represented a genuinely new capability curve or simply a way to buy benchmark performance with money, and how much weight any single benchmark could bear as a proxy for general reasoning.

Referenced by