Timeline

OpenAI publishes HealthBench, a medical-conversation benchmark

Built with 262 physicians across 60 countries, the open-source benchmark grades 5,000 simulated health conversations against physician-written rubrics.

  • Benchmarks & progress
  • Minor

OpenAI published HealthBench, an open-source benchmark for evaluating how language models perform in realistic health conversations, built with 262 physicians who had practised across 60 countries and 49 languages, spanning 26 medical specialties.

The benchmark consists of 5,000 simulated conversations between a model and either a layperson or a clinician, covering scenarios such as emergencies, handling uncertainty and global health contexts. Each conversation carries a custom rubric written by the physicians involved, breaking the ideal response into specific, gradable criteria — around 48,000 individual criteria in total — rather than scoring against a single reference answer. OpenAI said this rubric-based design was meant to capture the open-endedness of real medical dialogue more faithfully than prior multiple-choice medical benchmarks such as MedQA.

The release came as OpenAI and other labs pushed further into healthcare applications, and followed criticism that earlier medical benchmarks had become saturated by models scoring near the ceiling without corresponding gains in real clinical usefulness. HealthBench was released with an open licence so other developers and researchers could use it to evaluate their own models, and OpenAI said it planned to keep the benchmark updated as models improved.