Timeline

Google DeepMind proposes a 'Levels of AGI' framework

Six tiers from 'no AI' to 'superhuman', scored across narrow and general tasks, aimed to replace binary AGI-or-not debate with a shared measurement vocabulary.

  • Ideas & essays
  • Benchmarks & progress
  • Notable

A team of Google DeepMind researchers including Meredith Ringel Morris, Shane Legg and Allan Dafoe published a paper proposing a graded framework for classifying progress toward artificial general intelligence, arguing that the field lacked a shared, operational vocabulary for what “AGI” meant and was consequently arguing past itself whenever the term came up.

The framework scored systems along two axes: depth, meaning performance level, and breadth, meaning how general the capability was — a narrow system excelling at one task versus a general one competent across many. Performance was divided into six tiers: No AI, Emerging (matching or slightly exceeding an unskilled human), Competent (performance around the 50th percentile of skilled adults), Expert (90th percentile), Virtuoso (99th percentile), and Superhuman (exceeding all humans). The authors argued current systems such as ChatGPT and Bard already qualified as “Emerging AGI” — general-purpose but not yet reliably competent — under their own scheme, a claim intended to demonstrate the framework’s use rather than to assert any conclusion about how close true AGI was.

The paper’s stated purpose was practical rather than philosophical: to give researchers, companies and policymakers a common language for comparing systems, assessing risk and measuring progress, replacing binary “is this AGI or not” arguments — of the kind Microsoft Research’s “Sparks of AGI” paper had provoked earlier that year — with a graded scale that acknowledged capability accrues unevenly across different tasks and domains.

The framework was widely cited in subsequent debates over AI capability classification, though it did not settle the underlying disagreement it was designed to defuse: critics continued to dispute where particular systems fell on the scale, and whether depth and breadth could be meaningfully separated for systems whose performance varied enormously by task. It became a standard reference point alongside benchmark suites for structuring arguments about capability progress through the following years, even as no consensus formed around adopting its specific tier labels in mainstream industry use.