Timeline

UC Berkeley releases Agents' Last Exam, a benchmark of professional work

Built with 250+ industry experts across 55 sub-industries, it runs agents in the real software a specialist would use and grades against hidden answers; current systems clear under 1% of the hardest tier.

  • Benchmarks & progress
  • Notable

A group at UC Berkeley’s Center for Responsible, Decentralized Intelligence, led in part by Dawn Song, released Agents’ Last Exam, a benchmark that measures whether AI agents can finish real professional work rather than answer exam-style questions. Built with more than 250 industry experts, its tasks span 55 sub-industries mapped onto the US government’s own occupational taxonomy — from Adobe After Effects animation to Siemens NX 3D modelling to neuroimaging analysis — and hand an agent the actual software and input files it would need.

Each task is graded automatically against a reference answer kept hidden until the run completes, closing the usual routes to gaming a fixed test set. The initial release covered over a thousand tasks, with the authors stating an intention to grow the pool toward five thousand and to keep it a “living” benchmark rather than a one-off snapshot. Results were stark by design: on the hardest difficulty tier, the paper reported an average full-pass rate below 1% for current agent systems.

The framing echoed a wider shift through 2025 and 2026, visible also in OpenAI’s GDPval, toward evaluations anchored to economically valuable output rather than academic subjects — an attempt to measure how close agents were to doing real jobs, at a point when older benchmarks were saturating and self-reported scores were increasingly hard to interpret.