METR proposes 'expenditure horizon' measure
The metric prices AI agents against human effort in dollars per unit of progress; on a public speed-optimisation task, frontier agents matched roughly $3,300 of skilled human labour.
- Benchmarks & progress
- Minor
METR proposed a new metric, “expenditure horizon,” for measuring how well an AI agent can optimise something relative to a human doing the same work. Rather than comparing task success rates, it plots two cost curves — return on human labour spent and return on agent spend, both denominated in dollars — and reads off the point where the two curves cross.
METR tested the approach against the NanoGPT speedrun, a public community challenge to minimise training time to a target loss on eight H100 GPUs, using 82 recorded human contributions submitted between May 2024 and April 2026 as the human baseline, estimated at roughly $2,500 per percentage point of efficiency gained. Frontier agents, including models from OpenAI and Anthropic, were then run against the same task at increasing expenditure. Their expenditure horizons ranged up to about $3,300 — the point beyond which additional agent spending stopped matching what the same money would buy in human effort — with total autonomous gains of only around 1–1.5% after runs costing over $10,000. Human maintainers judged only 50–70% of the agents’ contributions genuinely mergeable.
METR’s own conclusion was that “autonomous agent optimization has so far had minimal effect on AI R&D progress” on this particular task, and that human labour remained more cost-effective at scale. The result sits alongside METR’s separate time-horizon work as a second, cost-based way of tracking how close agentic AI is to substituting for skilled researchers rather than merely assisting them.