Timeline

Epoch AI reports GPT-5's FrontierMath performance

Running its own scaffold rather than OpenAI's, Epoch scored GPT-5 at 24.8% on FrontierMath's main tiers and 8.3% on the hardest tier, a new high for the benchmark.

  • Benchmarks & progress
  • Minor

Epoch AI, an independent research group that tracks AI capability trends, published its own evaluation of GPT-5 on FrontierMath, a benchmark of unpublished, expert-level mathematics problems that Epoch itself maintains. Running the model with high reasoning effort on Epoch’s own evaluation scaffold rather than one supplied by OpenAI, the group reported a score of 24.8% on the benchmark’s main three tiers and 8.3% on Tier 4, the hardest and most recently added set of problems — a new high score on the benchmark at the time.

The evaluation mattered as an outside check on launch-week capability claims. Model releases are typically accompanied by benchmark numbers chosen and run by the lab itself; Epoch’s FrontierMath scores were produced independently, using problems that were not part of GPT-5’s training data by design, and gave researchers a comparison point against OpenAI’s own o3 model, which had previously held the benchmark’s high score.

The result was read as incremental rather than a step change: GPT-5 extended an existing OpenAI lead on FrontierMath rather than closing a large gap held by a competitor, and even the improved score meant the model still failed roughly three-quarters of the benchmark’s main-tier problems and over 90% of Tier 4. Epoch continued to publish comparable evaluations for subsequent model releases, and FrontierMath scores became one of the standard reference points cited in coverage of frontier model launches through 2025 and 2026.