Benchmarks · Reasoning & problem-solving

GPQA

also: GPQA Diamond, Graduate-Level Google-Proof Q&A Benchmark

Whether a model can answer graduate-level science questions that a skilled non-expert cannot solve even with unrestricted web access and half an hour per question.

NYU, Cohere & Anthropic researchersReleased 20 November 2023Saturating

GPQA’s name is a description of what it was built to defeat: “Google-proof” questions, graduate-level enough that looking them up doesn’t help. Researchers published it in November 2023 with 448 multiple-choice questions in biology, physics and chemistry, each written and checked by a PhD holder in that exact specialism. The paper’s baselines made the point sharply: domain experts scored 65% on questions inside their own field, but skilled non-experts given unrestricted web access and more than half an hour per question managed only 34% — barely above chance. GPT-4, the best model tested, reached 39%.

The gap the benchmark was designed to expose — between people who could look things up and people who actually understood the material — turned out to be exactly the gap reasoning-focused models started closing fastest. Scores climbed quickly through 2024: Claude 3.5 Sonnet improved on its predecessor, and by December 2024 OpenAI’s o3 reported 87.7% on the harder 198-question “Diamond” subset that had become the standard citation — well past the original PhD-expert baseline.

That baseline stopped being much of a ceiling. By late 2025, Gemini 3 Pro reported 91.9% on GPQA Diamond, rising to 93.8% in its higher-effort Deep Think mode, and GPQA had joined the small set of benchmarks — alongside MMLU before it — that most frontier labs cite in every release even as scores approach saturation. It remains one of the most consistently reported science benchmarks precisely because its original difficulty took so long to fall.

The set

448 multiple-choice questions in biology, physics and chemistry, written and validated by PhD holders in the relevant specialism. The hardest 198-question subset, GPQA Diamond, is the version almost universally reported in model release announcements.

Example

A reaction of a liquid organic compound, which molecules consist of carbon and hydrogen atoms, is performed at 80 centigrade and 20 bar for 24 hours. In the proton nuclear magnetic resonance spectrum, the signals with the highest chemical shift of the reactant are replaced by a signal of the product that is observed about three to four units downfield. Compounds from which position in the periodic system of the elements, which are also used in the corresponding large-scale industrial process, have been mostly likely initially added in small amounts? A) A metal compound from the fifth period. B) A metal compound from the fifth period and a non-metal compound from the third period. C) A metal compound from the fourth period. D) A metal compound from the fourth period and a non-metal compound from the second period.arxiv.org

Where it stands

Frontier reasoning models now score above 90% on GPQA Diamond, well past the paper's own PhD-expert baseline of 65%, and labs have begun citing harder successors alongside it.

How the top score changed hands

  1. November 2023GPT-439%Strongest model baseline the original paper tested, against a 65% PhD-expert baseline and 34% for skilled non-experts with web access.
  2. June 2024Claude 3.5 Sonnetimproved over Claude 3 OpusAnthropic reported gains on GPQA without publishing a single headline figure in this announcement.
  3. December 2024OpenAI o387.7%On GPQA Diamond, first announced at early access.
  4. November 2025Gemini 3 Pro91.9% (93.8% Deep Think)

Current best: Gemini 3 Pro (Deep Think) — 93.8% 91.9% for the standard Gemini 3 Pro configuration; 93.8% for the higher-effort Deep Think mode. Later frontier models may score higher but are not yet documented here.

In the timeline · 17 entries · showing 16 most notable

More reasoning & problem-solving benchmarks