Timeline

OpenAI publishes SimpleQA, a benchmark for factuality

A short-form factuality benchmark designed to be more challenging and less saturated than prior QA benchmarks.

  • Benchmarks & progress
  • Minor

OpenAI released SimpleQA, a benchmark of 4,326 short, fact-seeking questions each with a single unambiguous, verified answer, designed to isolate factual recall from reasoning and to remain difficult for frontier models rather than being immediately saturated.

Each question was written to have exactly one correct short answer and checked by two independent reviewers, spanning domains including history, science, technology and entertainment. The design goal was narrower than most QA benchmarks: rather than testing reasoning or comprehension, it aimed to measure a model’s calibration and propensity to hallucinate when asked a plain factual question with a knowable answer, and to grade answers as correct, incorrect or “not attempted.”

SimpleQA arrived within weeks of MLE-bench, part of a cluster of narrowly scoped OpenAI benchmark releases in October 2024 aimed at specific model weaknesses rather than broad capability. It was quickly adopted across the industry as a standard reference for measuring parametric factual accuracy, in part because its single-answer format made scoring unambiguous in a way that open-ended QA sets were not.