Timeline

OpenAI o1 results published on ARC-AGI-Pub

o1-preview scored 21% on the public evaluation set, similar to Claude 3.5 Sonnet, but took roughly 70 hours to run 400 tasks against 30 minutes for either non-reasoning model.

  • Benchmarks & progress
  • Minor

The ARC Prize Foundation published independent test results for OpenAI’s newly released o1 models against the public ARC-AGI evaluation set, a benchmark of visual pattern-completion puzzles designed to be easy for humans and resistant to memorisation by language models.

o1-preview scored 21.2% on the public evaluation set and o1-mini scored 12.8%, against 9% for GPT-4o and roughly 21% for Claude 3.5 Sonnet — putting o1-preview’s headline reasoning gain over OpenAI’s own prior model roughly in line with, not clearly ahead of, a competitor that used no comparable test-time reasoning step. The improvement came at a steep cost in compute time: ARC Prize reported that evaluating o1-preview on 400 public tasks took around 70 hours, versus roughly 30 minutes for GPT-4o or Claude 3.5 Sonnet, an average of about 4.2 minutes per task against under 20 seconds for the faster models.

ARC Prize’s assessment was measured rather than triumphant. It judged o1 an evolution rather than a breakthrough, arguing the model still operated largely within the distribution of patterns present in its training data and synthetic chain-of-thought templates, and that scaling test-time compute — letting the model “think” longer before answering — improved scores on tasks amenable to extended search without resolving the underlying generalisation problem ARC-AGI was designed to probe. The organisation’s often-repeated conclusion, that new ideas rather than more compute were still needed to close the gap to human-level performance on the benchmark, made the o1 results a data point in the argument for and against reasoning-through-search as a path to more general capability, rather than a resolution of it.

Referenced by