Benchmarks · Reasoning & problem-solving

Humanity's Last Exam

also: HLE

Whether a model can answer the hardest closed-ended questions expert academics could write in their own field, at a difficulty chosen specifically to be far from saturated.

Center for AI Safety & Scale AIReleased 23 January 2025Live

Humanity’s Last Exam was built to solve a specific problem: MMLU and its peers had stopped being useful once frontier models cleared roughly 90% accuracy, leaving no way to see continued progress. CAIS and Scale AI assembled 2,500 questions submitted by nearly 1,000 academics across roughly 500 institutions, each one filtered to be unanswerable by internet search and gradable by a fixed answer key. At launch in January 2025, the design succeeded almost too well — every model tested, including OpenAI’s o1 and GPT-4o, scored under 10%, and CAIS director Dan Hendrycks said the team genuinely did not know how fast scores would climb.

They climbed fast. Reasoning-focused models pushed past 40% within the year: xAI reported 44.4% for a multi-agent configuration of Grok 4, though the figure had not yet cleared independent verification on the public leaderboard, and Gemini 3 Pro reached 37.5% without tools, rising to 41.0% in its higher-effort Deep Think mode. By February 2026, an upgraded Deep Think reported 48.4% without external tools.

Scores with tool access — letting a model search or use a code interpreter mid-answer — have run higher still, which makes cross-model comparison harder: a number is only meaningful alongside the conditions it was measured under. HLE has nonetheless become, alongside GPQA and ARC-AGI, one of the benchmarks labs reach for specifically because it has not yet saturated, and its scores remain a standard citation in frontier model announcements.

The set

2,500 questions submitted by nearly 1,000 contributors across roughly 500 institutions in about 50 countries, spanning mathematics, the humanities and the natural sciences in both text-only and multimodal form. Questions were filtered to exclude anything answerable by internet search and to remain automatically gradable.

Example

Hummingbirds within Apodiformes uniquely have a bilaterally paired oval bone, a sesamoid embedded in the caudolateral portion of the expanded, cruciate aponeurosis of insertion of m. depressor caudae. How many paired tendons are supported by this sesamoid bone? Answer with a number.lastexam.ai

Where it stands

Every model tested at launch in January 2025 scored under 10%; by early 2026 leading models with extended reasoning had reached the high 40s to low 50s, still well short of saturation.

How the top score changed hands

  1. January 2025o1 / GPT-4ounder 10%Every model CAIS and Scale AI tested at launch, across the board, scored below 10%.
  2. July 2025Grok 4 Heavy44.4%Multi-agent 'Heavy' tier; xAI's figure had not yet appeared on the public leaderboard for independent verification.
  3. November 2025Gemini 3 Pro37.5% (41.0% Deep Think)Without external tools.
  4. February 2026Gemini 3 Deep Think (v2)48.4%

Current best: Gemini 3 Deep Think (v2) — 48.4% Without external tools. Moonshot AI separately reported 54.0% for Kimi K2.6 with tool use in April 2026 — a different, tool-assisted condition — in vendor comparison tables the company itself said were not independently verified.

In the timeline · 12 entries

More reasoning & problem-solving benchmarks