Timeline

GPQA graduate-level 'Google-proof' benchmark published

PhD-level experts scored 65% and skilled non-experts with unrestricted web access for over 30 minutes managed only 34%; GPT-4 reached 39%.

  • Benchmarks & progress
  • Notable

Researchers published GPQA, a set of 448 multiple-choice questions in biology, physics and chemistry written by domain experts and designed to resist being answered by looking things up. The benchmark’s name — “Google-proof” — described its intended property directly: questions hard enough that search access alone would not close the gap between an expert and a non-expert.

The paper reported that PhD-level experts, tested on questions inside their own specialism, reached 65% accuracy (74% if their own subsequently identified errors were excluded), while highly skilled non-experts given unrestricted web access and more than 30 minutes per question managed only 34% — barely above the 25% expected from guessing across four options. GPT-4, the strongest model baseline the authors tested, scored 39%, ahead of the non-expert humans but far short of the domain experts.

The benchmark’s stated purpose was narrower than a general capability leaderboard: it was built to support “scalable oversight” research, the problem of how humans might usefully supervise or evaluate an AI system whose answers they cannot independently verify. A benchmark where non-experts with search access could not close the gap to domain experts gave researchers a controlled setting to test methods — such as debate between models, or AI-assisted human judgment — for closing that gap some other way.

GPQA went on to become one of the standard benchmarks frontier labs reported alongside each new model release through 2024 and into 2025, valued precisely because its difficulty resisted the search-and-lookup strategies that had eroded other knowledge tests. Later reasoning-focused models pushed scores well past the original human-expert baseline, and by 2025 the benchmark was widely treated as approaching saturation, prompting labs to adopt harder successors.