Benchmarks · Reasoning & problem-solving
GPQA
also: GPQA Diamond, Graduate-Level Google-Proof Q&A Benchmark
Whether a model can answer graduate-level science questions that a skilled non-expert cannot solve even with unrestricted web access and half an hour per question.
NYU, Cohere & Anthropic researchersReleased 20 November 2023
GPQA’s name is a description of what it was built to defeat: “Google-proof” questions, graduate-level enough that looking them up doesn’t help. Researchers published it in November 2023 with 448 multiple-choice questions in biology, physics and chemistry, each written and checked by a PhD holder in that exact specialism. The paper’s baselines made the point sharply: domain experts scored 65% on questions inside their own field, but skilled non-experts given unrestricted web access and more than half an hour per question managed only 34% — barely above chance. GPT-4, the best model tested, reached 39%.
The gap the benchmark was designed to expose — between people who could look things up and people who actually understood the material — turned out to be exactly the gap reasoning-focused models started closing fastest. Scores climbed quickly through 2024: Claude 3.5 Sonnet improved on its predecessor, and by December 2024 OpenAI’s o3 reported 87.7% on the harder 198-question “Diamond” subset that had become the standard citation — well past the original PhD-expert baseline.
That baseline stopped being much of a ceiling. By late 2025, Gemini 3 Pro reported 91.9% on GPQA Diamond, rising to 93.8% in its higher-effort Deep Think mode, and GPQA had joined the small set of benchmarks — alongside MMLU before it — that most frontier labs cite in every release even as scores approach saturation. It remains one of the most consistently reported science benchmarks precisely because its original difficulty took so long to fall.
The set
448 multiple-choice questions in biology, physics and chemistry, written and validated by PhD holders in the relevant specialism. The hardest 198-question subset, GPQA Diamond, is the version almost universally reported in model release announcements.
Example
A reaction of a liquid organic compound, which molecules consist of carbon and hydrogen atoms, is performed at 80 centigrade and 20 bar for 24 hours. In the proton nuclear magnetic resonance spectrum, the signals with the highest chemical shift of the reactant are replaced by a signal of the product that is observed about three to four units downfield. Compounds from which position in the periodic system of the elements, which are also used in the corresponding large-scale industrial process, have been mostly likely initially added in small amounts? A) A metal compound from the fifth period. B) A metal compound from the fifth period and a non-metal compound from the third period. C) A metal compound from the fourth period. D) A metal compound from the fourth period and a non-metal compound from the second period.arxiv.org
Where it stands
Frontier reasoning models now score above 90% on GPQA Diamond, well past the paper's own PhD-expert baseline of 65%, and labs have begun citing harder successors alongside it.
How the top score changed hands
- November 2023GPT-439%Strongest model baseline the original paper tested, against a 65% PhD-expert baseline and 34% for skilled non-experts with web access.
- June 2024Claude 3.5 Sonnetimproved over Claude 3 OpusAnthropic reported gains on GPQA without publishing a single headline figure in this announcement.
- December 2024OpenAI o387.7%On GPQA Diamond, first announced at early access.
- November 2025Gemini 3 Pro91.9% (93.8% Deep Think)
Current best: Gemini 3 Pro (Deep Think) — 93.8% 91.9% for the standard Gemini 3 Pro configuration; 93.8% for the higher-effort Deep Think mode. Later frontier models may score higher but are not yet documented here.
In the timeline · 17 entries · showing 16 most notable
Moonshot AI launches Kimi K3
A mixture-of-experts design activating 104 billion of its 2.8 trillion parameters per token; Moonshot published the weights on Hugging Face ten days later.
Open weights & ecosystem · Models & capabilities
Moonshot AI releases Kimi K2.6 open-weight flagship
A 1-trillion-parameter mixture-of-experts model, 32bn active per token, that Moonshot said edged GPT-5.4 on SWE-Bench Pro while costing several times less to run.
Open weights & ecosystem · Models & capabilities · Benchmarks & progress
Google ships Gemini 3
Gemini 3 Pro reported a 1501 Elo score on LMArena and 91.9% on GPQA Diamond, prompting OpenAI to reportedly declare an internal 'code red' days later.
Models & capabilities · Benchmarks & progress
Nous Research releases Hermes 4
Built by post-training Llama 3.1 checkpoints alone, the 405B model scored 57.1% on RefusalBench against 17.67% for GPT-4o, reflecting Nous's low-refusal alignment approach.
Open weights & ecosystem · Models & capabilities
METR examines how time horizon varies across domains
Applying its 50%-success task-length method to nine benchmarks, METR found doubling times of two to six months for reasoning tasks but around twenty months for Tesla's self-driving system.
Benchmarks & progress
Google updates Gemini 2.5 Pro preview with improved coding performance
The update, internally labelled 06-05, also led coding benchmarks including Aider Polyglot and performed strongly on Humanity's Last Exam.
Models & capabilities · Benchmarks & progress
DeepSeek releases DeepSeek-R1-0528 update
Released under an MIT licence, the update raised AIME 2025 accuracy from 70% to 87.5% by roughly doubling the average length of the model's reasoning traces.
Open weights & ecosystem · Models & capabilities · Benchmarks & progress
Stanford HAI releases 2025 AI Index Report
The eighth annual report put US private AI investment at $109.1 billion in 2024, nearly twelve times China's $9.3 billion, and inference cost for GPT-3.5-level performance down over 280-fold since late 2022.
Benchmarks & progress · Ideas & essays
Gemini 2.5 Pro takes the lead on reasoning benchmarks
Google's thinking model topped LMArena and several reasoning evaluations, its strongest competitive position of the period.
Models & capabilities · Benchmarks & progress
DeepSeek releases DeepSeek-V3-0324 update
The updated checkpoint scored 81.2% on MMLU-Pro and 59.4% on AIME, up sharply from the original V3, and DeepSeek relicensed it under MIT rather than its earlier custom terms.
Open weights & ecosystem · Models & capabilities
OpenAI releases GPT-4.5
Priced at $75/$150 per million tokens, about thirty times GPT-4o's rate, and retired from the API within five months in favour of the cheaper GPT-4.1.
Models & capabilities
xAI releases Grok-3
xAI reported Grok 3 beating GPT-4o and o3-mini-high on AIME and GPQA using roughly ten times the compute of Grok 2, on figures the company had not independently verified.
Models & capabilities · Benchmarks & progress
OpenAI announces o3 and opens early access for safety testing
Reported scores included 96.7% on the AIME maths exam and a Codeforces rating in the 99.2nd percentile; OpenAI cited o1's link between reasoning and deception as a reason to delay release.
Models & capabilities · Benchmarks & progress
xAI releases Grok-2
The beta release added image generation via Black Forest Labs' FLUX.1 and, within days, took second place on the LMSYS Chatbot Arena leaderboard behind GPT-4o.
Models & capabilities
Claude 3.5 Sonnet and Artifacts change how people use chatbots
Priced and sped like Anthropic's mid-tier model, it scored 64% on the company's internal agentic-coding evaluation against 38% for the outgoing flagship.
Models & capabilities
GPQA graduate-level 'Google-proof' benchmark published
PhD-level experts scored 65% and skilled non-experts with unrestricted web access for over 30 minutes managed only 34%; GPT-4 reached 39%.
Benchmarks & progress