Timeline

Gemini 2.5 Pro takes the lead on reasoning benchmarks

Google's thinking model topped LMArena and several reasoning evaluations, its strongest competitive position of the period.

  • Models & capabilities
  • Benchmarks & progress
  • Major

Google DeepMind released Gemini 2.5 Pro, an experimental model built on a “thinking” architecture that reasons through intermediate steps before producing an answer, rather than responding directly. It was made available in Google AI Studio and, for Gemini Advanced subscribers, in the Gemini app.

Google’s own announcement claimed the model topped the LMArena leaderboard, which ranks models by aggregated human preference in head-to-head comparisons, “by a significant margin” on release — a strong result for a company that had spent much of the prior two years trailing OpenAI and Anthropic on public leaderboards. The company also reported leading scores on the GPQA science-reasoning benchmark and the 2025 American Invitational Mathematics Examination without relying on the more expensive test-time techniques, such as running many samples and voting on the best, that some rivals used to boost scores. On Humanity’s Last Exam, a benchmark designed to be difficult even for frontier models, Gemini 2.5 Pro scored 18.8% without external tools; on SWE-bench Verified, a coding benchmark built from real GitHub issues, it reached 63.8% using a custom agent setup. The model also carried a one-million-token context window, which Google said would expand to two million.

Because benchmark leadership in this period tended to last only weeks before a competitor released something stronger, the significance of the release lay less in any individual score than in its timing: it arrived as Google, OpenAI and Anthropic were converging on the same architectural idea — giving models extended, visible or semi-visible reasoning steps before an answer — and Gemini 2.5 Pro was, for a period, the clearest public evidence that Google’s version of that approach was competitive with or ahead of its rivals’.