Benchmarks · Mathematics

Omni-MATH

Can a model solve genuinely Olympiad-level mathematics problems, of the kind that saturated benchmarks like MATH no longer contain?

Peking University & AlibabaReleased 10 October 2024Live

Omni-MATH exists because its two predecessors stopped working. By late 2024, GSM8K and MATH were both reporting frontier scores in the 90s, and researchers from Peking University and Alibaba built a replacement drawn from genuine mathematical Olympiad competitions rather than high-school contests. Its 4,428 problems span more than 33 sub-domains and ten difficulty tiers, aiming to separate models the way MATH once did before its ceiling was reached.

The gap it exposed was immediate. At publication, OpenAI’s reasoning models o1-preview and o1-mini scored 52.55% and 60.54% respectively — the clear best results in the paper — while GPT-4o, a strong non-reasoning model from the same period, managed roughly 30%. Contemporary open models built specifically for mathematics, such as Qwen2.5-MATH, trailed further still. The spread showed that extended reasoning at inference time, the technique behind OpenAI’s o1 line, mattered more for Olympiad-level problems than raw model scale alone.

Omni-MATH’s own published leaderboard has not visibly grown since shortly after the paper’s release, so later frontier scores from 2025 and 2026 models are not confirmed here even though they are widely reported to be substantially higher. The benchmark’s broader role — an Olympiad-difficulty stand-in once school-competition mathematics stopped being hard enough — sits alongside harder, less saturated successors such as FrontierMath in tracking how far mathematical reasoning has moved past what MATH could measure.

The set

4,428 competition-level problems with rigorous human annotation, organised into more than 33 mathematical sub-domains and over 10 difficulty tiers, drawn from Olympiad-tier competitions rather than the high-school-competition tier that MATH uses.

Example

The dataset's first problem, verbatim: 'Let n(≥2) be a positive integer. Find the minimum m, so that there exists x_ij (1≤i,j≤n) satisfying...' three conditions on running row and column maxima. Answer: 1+⌈n/2⌉ — an Olympiad-tier combinatorics problem, not the school-competition tier MATH used.huggingface.co

Where it stands

Built explicitly because GSM8K and MATH had stopped separating frontier models; the project's own leaderboard has not been updated since late 2024, so later frontier scores are not independently confirmed here.

How the top score changed hands

  1. May 2024GPT-4o30.49%For comparison, a non-reasoning frontier model of the same period scored roughly half of o1-mini's result.
  2. September 2024OpenAI o1-preview52.55%Reported in the original paper alongside o1-mini as the two standout results among tested models.

Current best: OpenAI o1-mini — 60.54% The top score on the benchmark's own leaderboard as published; later reasoning models are reported to score higher but are not confirmed against this leaderboard.

More mathematics benchmarks