Timeline

OpenAI publishes ten formally-verified math advances from unreleased Astra model

OpenAI said generating all ten proofs cost about $2,000 in compute; mathematician Gary Marcus called the framing 'vastly oversold' relative to what the paper actually verified.

  • Benchmarks & progress
  • Models & capabilities
  • Ideas & essays
  • Major

OpenAI published results from an internal version of a previously undisclosed model, Astra, which it described as its next major model family, alongside ten machine-checkable solutions to open problems spanning group theory, coding theory, quantum complexity, lattice cryptography and extremal combinatorics. Among them was a construction establishing the existence of non-sofic groups, addressing a question open since the 1990s, and a disproof of a rigidity conjecture associated with Alain Connes. OpenAI said generating all ten solutions cost roughly $2,000 at its internal API rates and said each proof had been formalised in Lean, producing a machine-checkable certificate rather than relying on the model’s own claim of correctness.

Reaction split along familiar lines. Mathematician Thomas Bloom of the University of Manchester called the results “big news.” Cognitive scientist Gary Marcus, a longstanding critic of claims made for large language models, argued the announcement was “vastly oversold”: he noted that OpenAI’s accompanying paper said little about how the model worked, what role human mathematicians played in framing or checking the problems, or whether any of the model’s other proposed proofs had failed, and he and mathematician Ernie Davis raised the possibility that OpenAI had selected favourable results from a larger, unpublished set of attempts. Marcus also argued that mathematics was unusually well suited to this kind of demonstration because proofs can be checked symbolically and training data can be generated synthetically with guaranteed correctness, a combination that does not carry over to most open-ended real-world tasks.

The announcement was OpenAI’s first public confirmation of the Astra name, positioned as a successor to its GPT-5.6 generation, though the company had not said whether it would ship under that name or as a numbered GPT release. As with earlier OpenAI mathematics claims, the results had not been peer-reviewed at the time of publication, and the dispute over how much of the achievement reflected genuine domain generality — as opposed to mathematics-specific advantages in verification and data generation — echoed arguments that followed nearly every prior benchmark claim in the field.