Timeline

METR publishes 'Measuring AI Ability to Complete Long Software Tasks'

Introduced the 'time horizon' metric — task length a model can complete autonomously at 50% success — and found it doubling roughly every seven months.

  • Ideas & essays
  • Benchmarks & progress
  • Major

METR, a nonprofit that evaluates frontier models for dangerous capabilities, published a paper defining a new way to measure AI progress: the “50%-task-completion time horizon,” the length of a task — measured in the time a skilled human would need — that a model can complete autonomously with 50% success. The measure was built from two existing benchmark suites, RE-Bench and HCAST, plus 66 new shorter tasks the authors added to fill out the low end of the difficulty range, mostly software and reasoning problems.

Fitting the results to human baseline times, the paper reported that the frontier model at the time, Claude 3.7 Sonnet, had a time horizon of roughly 50 minutes — succeeding close to 100% of the time on tasks under about four minutes, but less than 10% of the time on tasks over roughly four hours. Tracing the same measure back through earlier model generations, the authors found the horizon had been doubling approximately every seven months since 2019, with some evidence of acceleration in the most recent period. Extrapolated forward, the trend implied that within a few years models could complete tasks that currently take a skilled person a month.

The authors were explicit about the uncertainty involved: task selection, the choice of human comparison cohort and methodological judgement calls could shift the estimated horizon by an order of magnitude, which alone would move a forecast arrival date by around two years, and the paper’s own extrapolation assumed the trend would continue rather than plateau or bend.

The metric was taken up quickly as a reference point in debates about AI timelines. It became a standard citation in arguments over how close autonomous AI research or lengthy agentic work might be, including in the AI Futures Project’s “AI 2027” scenario published two weeks later, and in the critiques that followed it.