Epoch AI expands FrontierMath to 50 unsolved research problems
Unlike FrontierMath's original tiers, these problems have no known solution at all; three of the fifty have been solved by AI, including one by GPT-5.6 Sol.
- Benchmarks & progress
- Notable
Epoch AI expanded FrontierMath’s “Open Problems” tier from 14 to 50 questions, each drawn from genuine mathematical research that professional mathematicians have attempted and failed to resolve. It is a different kind of benchmark from FrontierMath’s original four tiers, where every problem has a known solution that is withheld from public view and used to grade a model’s answer; here there is no known answer to check against, so each problem instead ships with a bespoke verifier program that can confirm whether a proposed solution is correct without anyone knowing in advance what a correct solution looks like.
Three of the fifty problems have been solved since the benchmark launched. Epoch AI said OpenAI’s GPT-5.6 Sol beat the best-known upper bounds for superpermutations — a long-studied combinatorics problem — over sequences of 8, 9 and 10 symbols during an interactive session, one of three solved problems all classed in the benchmark’s “Moderately Interesting” difficulty tier.
The tier is a small-scale but concrete test of a claim frontier labs increasingly make: that their models can extend, not just recall, mathematical knowledge. Because each solved problem is independently verified rather than self-reported, and because the problems were unsolved by humans before an AI system solved them, Open Problems is one of the few benchmarks whose results are difficult to dismiss as data contamination or overfitting to a known answer.