Epoch AI launches FrontierMath
Built with over 60 mathematicians including Fields medallists as reviewers, the benchmark held leading models under 2% accuracy even with extended reasoning time and code tools.
- Benchmarks & progress
- Notable
Epoch AI launched FrontierMath, a benchmark of several hundred original, expert-crafted mathematics problems spanning fields including computational number theory, real analysis, algebraic geometry and category theory. The problems were developed and peer-reviewed by a network of more than 60 mathematicians, including professors, International Mathematical Olympiad question-setters, and Fields medallists Terence Tao and Timothy Gowers, both of whom were quoted calling the problems extremely challenging and, in Tao’s words, of a different order of difficulty from Olympiad questions.
The benchmark’s design deliberately targeted a weakness of existing maths evaluations, which large language models had increasingly saturated: FrontierMath problems were original, unpublished, and resistant to being solved by pattern-matching against training data. Epoch tested six leading models, including Claude 3.5 Sonnet, o1-preview, GPT-4o and Gemini 1.5 Pro, giving them extended reasoning time and access to a Python environment. All six solved fewer than 2% of problems, a sharp contrast with benchmarks such as MMLU or GSM8K where top models scored above 90%.
FrontierMath was positioned as a longer-lived proxy for advanced mathematical reasoning than benchmarks that models had already begun to saturate, and it was cited within weeks by OpenAI as evidence of o3’s step-change in reasoning performance. That prominence made a subsequent disclosure more consequential: two months later it emerged that OpenAI had funded FrontierMath’s development and held privileged access to its problems and solutions, a relationship Epoch had not disclosed to the mathematicians who built it, which raised questions about how independent the benchmark’s headline results actually were.