Timeline

ARC Prize introduces public ARC-AGI leaderboard

Unlike the private-evaluation Kaggle competition, the leaderboard allows internet access and unlimited compute; early verified scores ranged from 42% down to 8-9% for frontier chatbots.

  • Benchmarks & progress
  • Minor

The ARC Prize Foundation opened a standing public leaderboard for submissions to ARC-AGI, the abstract-reasoning benchmark designed by François Chollet to resist memorisation, complementing the Kaggle competition it had launched earlier that month. Where the competition restricted entrants to a private evaluation set with no internet access and limited compute, the public leaderboard let anyone test against the public evaluation tasks under looser conditions, with a $150,000 verification fund set aside to support entrants whose results needed independent checking.

Initial verified scores illustrated a wide gap between specialised and general-purpose approaches: independent researcher Ryan Greenblatt topped the board at 42%, ahead of a long-standing brute-force programme called “icecuber 2020” at 39%. Frontier chatbots scored far lower without task-specific engineering — Claude 3.5 reached 21%, GPT-4o 9% and Gemini 1.5 8% — underscoring the benchmark’s original purpose of separating pattern-matching from the kind of novel-task reasoning it was designed to probe.

The leaderboard gave ARC-AGI a continuous, independently verifiable public record alongside the annual competition, and became a reference point through 2024 and 2025 as labs published increasingly aggressive scores against the same tasks, culminating in disputes over how much of the improvement reflected genuine reasoning gains versus compute-heavy search.