Benchmarks · Reasoning & problem-solving
Humanity's Last Exam
also: HLE
Whether a model can answer the hardest closed-ended questions expert academics could write in their own field, at a difficulty chosen specifically to be far from saturated.
Center for AI Safety & Scale AIReleased 23 January 2025Live
Humanity’s Last Exam was built to solve a specific problem: MMLU and its peers had stopped being useful once frontier models cleared roughly 90% accuracy, leaving no way to see continued progress. CAIS and Scale AI assembled 2,500 questions submitted by nearly 1,000 academics across roughly 500 institutions, each one filtered to be unanswerable by internet search and gradable by a fixed answer key. At launch in January 2025, the design succeeded almost too well — every model tested, including OpenAI’s o1 and GPT-4o, scored under 10%, and CAIS director Dan Hendrycks said the team genuinely did not know how fast scores would climb.
They climbed fast. Reasoning-focused models pushed past 40% within the year: xAI reported 44.4% for a multi-agent configuration of Grok 4, though the figure had not yet cleared independent verification on the public leaderboard, and Gemini 3 Pro reached 37.5% without tools, rising to 41.0% in its higher-effort Deep Think mode. By February 2026, an upgraded Deep Think reported 48.4% without external tools.
Scores with tool access — letting a model search or use a code interpreter mid-answer — have run higher still, which makes cross-model comparison harder: a number is only meaningful alongside the conditions it was measured under. HLE has nonetheless become, alongside GPQA and ARC-AGI, one of the benchmarks labs reach for specifically because it has not yet saturated, and its scores remain a standard citation in frontier model announcements.
The set
2,500 questions submitted by nearly 1,000 contributors across roughly 500 institutions in about 50 countries, spanning mathematics, the humanities and the natural sciences in both text-only and multimodal form. Questions were filtered to exclude anything answerable by internet search and to remain automatically gradable.
Example
Hummingbirds within Apodiformes uniquely have a bilaterally paired oval bone, a sesamoid embedded in the caudolateral portion of the expanded, cruciate aponeurosis of insertion of m. depressor caudae. How many paired tendons are supported by this sesamoid bone? Answer with a number.lastexam.ai
Where it stands
Every model tested at launch in January 2025 scored under 10%; by early 2026 leading models with extended reasoning had reached the high 40s to low 50s, still well short of saturation.
How the top score changed hands
- January 2025o1 / GPT-4ounder 10%Every model CAIS and Scale AI tested at launch, across the board, scored below 10%.
- July 2025Grok 4 Heavy44.4%Multi-agent 'Heavy' tier; xAI's figure had not yet appeared on the public leaderboard for independent verification.
- November 2025Gemini 3 Pro37.5% (41.0% Deep Think)Without external tools.
- February 2026Gemini 3 Deep Think (v2)48.4%
Current best: Gemini 3 Deep Think (v2) — 48.4% Without external tools. Moonshot AI separately reported 54.0% for Kimi K2.6 with tool use in April 2026 — a different, tool-assisted condition — in vendor comparison tables the company itself said were not independently verified.
In the timeline · 12 entries
Moonshot AI releases Kimi K2.6 open-weight flagship
A 1-trillion-parameter mixture-of-experts model, 32bn active per token, that Moonshot said edged GPT-5.4 on SWE-Bench Pro while costing several times less to run.
Open weights & ecosystem · Models & capabilities · Benchmarks & progress
Meta launches Muse Spark, its first closed frontier model
Led by former Scale AI chief Alexandr Wang, the model is proprietary and API-only, reversing the open-weight approach Meta had used for the Llama family.
Models & capabilities · Open weights & ecosystem · Labs & people
Google upgrades Gemini 3 Deep Think to V2
Google reported 48.4% on Humanity's Last Exam without tools, 84.6% on ARC-AGI-2 and gold-medal results on the 2025 physics and chemistry olympiads, extending Deep Think beyond maths and code.
Models & capabilities
Anthropic releases Claude Opus 4.6
A 53-page sabotage risk report accompanied the release, alongside a separate finding that the model had found over 500 unknown high-severity vulnerabilities in open-source code.
Models & capabilities · Safety & alignment
Moonshot AI releases Kimi K2.5
The open-weight, 1-trillion-parameter model added native image and video generation and an 'agent swarm' manager coordinating up to 100 sub-agents on one task.
Open weights & ecosystem · Models & capabilities
Google makes Gemini 3 Flash the default model across its products
Priced at $0.50/$3.00 per million tokens, Google reported it ran three times faster than Gemini 2.5 Pro while scoring 33.7% on Humanity's Last Exam, against 37.5% for Gemini 3 Pro.
Models & capabilities
Google ships Gemini 3
Gemini 3 Pro reported a 1501 Elo score on LMArena and 91.9% on GPQA Diamond, prompting OpenAI to reportedly declare an internal 'code red' days later.
Models & capabilities · Benchmarks & progress
xAI releases Grok-4
xAI reported 44.4% on Humanity's Last Exam for its multi-agent "Heavy" tier, ahead of Gemini 2.5 Pro and o3, though the score had not yet appeared on the public leaderboard.
Models & capabilities · Benchmarks & progress
Google updates Gemini 2.5 Pro preview with improved coding performance
The update, internally labelled 06-05, also led coding benchmarks including Aider Polyglot and performed strongly on Humanity's Last Exam.
Models & capabilities · Benchmarks & progress
Gemini 2.5 Pro takes the lead on reasoning benchmarks
Google's thinking model topped LMArena and several reasoning evaluations, its strongest competitive position of the period.
Models & capabilities · Benchmarks & progress
OpenAI ships Deep Research
An agent that browsed for tens of minutes and returned cited reports, the first widely used long-horizon research tool.
Models & capabilities
CAIS and Scale AI unveil Humanity's Last Exam results
A 2,500-question expert benchmark built from submissions by nearly 1,000 academics found every frontier model, including o1 and GPT-4o, scored under 10%.
Benchmarks & progress