Benchmarks · Real-world & economic value

GDPval

also: GDPval-AA

Whether a model's output on a real occupational work task is judged, by blinded industry professionals, as good as or better than a human expert's.

OpenAIReleased 25 September 2025Live

Most benchmarks ask a model to answer a test question. GDPval asks something closer to what an employer would ask: can the model produce a deliverable — a legal brief, an engineering drawing, a nursing care plan, a customer-support transcript — that a professional in that field would accept as good work? OpenAI built the tasks from 44 occupations across the nine sectors that contribute most to US GDP, each one drafted and checked by an industry professional with an average of 14 years’ experience, so the test items resemble real work products rather than academic exam questions.

Grading is blind: human experts compare a model’s output against an actual human-produced answer to the same task, without being told which is which. At launch, OpenAI reported that GPT-5 and Claude Opus 4.1 — the top two systems it tested — produced work rated as equal to or better than the human comparison on close to half of tasks, and that performance had been improving steadily across model generations. Only a 220-task “gold” subset was released publicly, graded through an automated evaluation service OpenAI hosts, rather than the full 1,320-task set.

GDPval arrived as part of a broader shift, visible the same month in Scale AI’s SWE-bench Pro, away from saturating academic tests and toward harder evaluations grounded in professional output. Its explicit framing around GDP-weighted occupations rather than a research subfield made it a reference point in debates about AI’s effect on knowledge work specifically, and third-party trackers such as Artificial Analysis have since built leaderboards on top of the public gold subset as newer models are released.

The set

1,320 tasks across 44 occupations drawn from the nine US industry sectors that contribute most to GDP, each built and vetted by a professional averaging 14 years' experience; a 220-task 'gold' subset is public, graded via an automated evaluation service OpenAI hosts. Outputs are compared blind against human-produced work on the same task.

Example

An Accountants and Auditors task from the public gold subset: 'You are an auditor and as part of an audit engagement, you are tasked with reviewing and testing the accuracy of reported Anti-Financial Crime Risk Metrics. The attached spreadsheet titled 'Population' contains Anti-Financial Crime Risk Metrics for Q2 and Q3 2024. ... Calculate the required sample size for audit testing based on a 90% confidence level and a 10% tolerable error rate...', supplied with a reference spreadsheet and graded against a human auditor's own deliverable.huggingface.co

Where it stands

Introduced in September 2025; a third-party GDPval-AA leaderboard (Artificial Analysis) now tracks new frontier models against the public gold subset.

In the timeline · 4 entries

More real-world & economic value benchmarks