Threads

Benchmarks and their saturation

The problem that AI benchmarks keep being saturated as fast as they are built — from GPT-4 topping professional exams to models earning perfect Olympiad marks — and the contamination that clouds the scores.

A benchmark is a test meant to measure progress; this thread follows what happens when the models keep passing them. GPT-4 set the pattern by scoring near the top of the human range on professional and academic exams, and the frontier models that answered it — Gemini, Claude 3 — competed largely on the same leaderboards.

Saturation forced harder tests. As standard benchmarks topped out, attention moved to ones designed to resist memorisation: OpenAI’s o1, and then o3’s breakthrough score on ARC-AGI — a test built specifically to be easy for humans and hard for models — marked how quickly even purpose-built evaluations fell. The scores also invited scrutiny: Llama 4’s troubled launch included accusations that a benchmark-tuned variant had been used to flatter its standing on a public leaderboard.

By 2025 the field had moved toward competition mathematics, where problems are fresh each year. Reasoning systems reached gold-medal standard at the IMO; in 2026 they scored perfect marks and began contributing formally-verified mathematical results. The thread’s running tension is measurement itself: when a model can be trained on, or contaminated by, the very test meant to judge it, a saturated benchmark says as much about the ruler as about the thing measured.