Benchmarks · Mathematics

FrontierMath

Can a model solve original, unpublished research-level mathematics problems that resist pattern-matching against training data?

Epoch AIReleased 8 November 2024Disputed

FrontierMath asks whether a model can do mathematics no one has fed it before: Epoch AI built several hundred original research-level problems, spanning fields from computational number theory to category theory, developed and peer-reviewed by more than 60 mathematicians including Fields medallists Terence Tao and Timothy Gowers. Because the problems were unpublished and resistant to pattern-matching, the benchmark was designed to outlast the fate of tests like MMLU, which models had already saturated. At launch, six leading models — including Claude 3.5 Sonnet, o1-preview and GPT-4o — all solved fewer than 2% of problems, even with extended reasoning time and a code interpreter.

Scores climbed as reasoning models improved: OpenAI cited FrontierMath within weeks as evidence of o3’s step-change in mathematical reasoning, and by August 2025 Epoch’s independent evaluation of GPT-5 put it at 24.8% on the main tiers and 8.3% on the hardest, most recently added Tier 4 — a new high, though still a majority failure rate. That prominence also drew scrutiny: two months after launch it emerged that OpenAI had funded FrontierMath’s development and held privileged access to its problems and solutions, a relationship Epoch had not disclosed to the mathematicians who built it.

The benchmark’s reliability was further complicated in May 2026, when Epoch disclosed that an AI-assisted audit had flagged fatal errors — mostly off-by-one slips and flipped signs in the answer key — in roughly a third of Tiers 1-4 problems, a figure that grew to 42% after full human review; a corrected v2 followed. The same year Epoch expanded a separate “Open Problems” tier to 50 genuinely unsolved research questions, graded by bespoke verifiers rather than a withheld answer, after AI systems had already solved three of the original fourteen — including OpenAI’s GPT-5.6 Sol beating known bounds on a long-studied combinatorics problem.

The set

Several hundred original problems across Tiers 1-4, spanning computational number theory, real analysis, algebraic geometry and category theory, developed and peer-reviewed by a network of over 60 mathematicians including Fields medallists; a separate Open Problems tier (expanded to 50 questions in mid-2026) uses genuine unsolved research problems graded by a bespoke verifier rather than a withheld answer.

Example

Construct a degree 19 polynomial p(x) in C[x] such that X := {p(x) = p(y)} in P^1 x P^1 has at least 3 (but not all linear) irreducible components over C. Choose p(x) to be odd, monic, have real coefficients and linear coefficient -19 and calculate p(19).epoch.ai

Where it stands

GPT-5 held the highest verified score at 24.8% on the main tiers as of August 2025; in May 2026 Epoch disclosed an AI-assisted audit had found fatal errors in roughly a third of Tiers 1-4 problems (later revised to 42% after full human review), and published a corrected v2 the following month.

How the top score changed hands

  1. November 2024Six frontier models (Claude 3.5 Sonnet, o1-preview, GPT-4o, Gemini 1.5 Pro and others)<2%All six models tested at launch, with extended reasoning time and a Python environment, solved fewer than 2% of problems — a sharp contrast with MMLU or GSM8K, where the same models scored above 90%.
  2. August 2025GPT-524.8% (Tiers 1-3), 8.3% (Tier 4)Independent Epoch evaluation, extending OpenAI's earlier lead on the benchmark rather than closing a gap held by a rival lab.

Current best: GPT-5 — 24.8% (Tiers 1-3), 8.3% (Tier 4) Scored by Epoch on its own evaluation scaffold rather than OpenAI's, with high reasoning effort; a new high at the time, though the model still failed roughly three-quarters of main-tier problems.

In the timeline · 6 entries

More mathematics benchmarks