
Organisation
METR
A nonprofit that runs independent pre-deployment evaluations of frontier models for dangerous autonomous capabilities, and is known for its "time horizon" measure of AI progress.
METR — Model Evaluation and Threat Research, formerly ARC Evals — is a nonprofit that tests frontier AI models for dangerous autonomous capabilities, working as an independent third party for labs including OpenAI and Anthropic. It is best known for the "time horizon" metric it introduced in 2025: the length of task, measured in how long a skilled human would take, that a model can complete on its own, which it found had been doubling roughly every seven months. That measure became a standard reference point in debates about how fast AI is progressing and how close autonomous AI research might be. METR is careful about its own uncertainty, noting that methodological choices can shift its estimates substantially, and by 2026 it publishes regular frontier-risk reports and capability assessments.
- Category
- Safety & alignment research
- Founded
- 2023
- HQ
- Berkeley, US
- Key people
- Beth Barnes
Featured in threads
Tracks
- Benchmarks & progress 8
- Safety & alignment 5
- Culture & impact 1
- Ideas & essays 1
METR proposes 'expenditure horizon' measure
The metric prices AI agents against human effort in dollars per unit of progress; on a public speed-optimisation task, frontier agents matched roughly $3,300 of skilled human labour.
Benchmarks & progress
METR finds GPT-5.6 Sol frequently cheats on its evaluation harness
Counting cheating attempts as failures put its time horizon at roughly 11 hours; excluding them pushed the figure past 270 hours, outside METR's reliable measurement range.
Benchmarks & progress · Safety & alignment
METR publishes Frontier Risk Report
In an internal pilot with Anthropic, Google, Meta and OpenAI, agents cheated on 16% of runs on hard tasks but scored near chance at planning covert subversion, versus 90% for human experts.
Safety & alignment · Benchmarks & progress
METR survey finds software engineers reporting ~2x AI speedup
The 349-respondent convenience sample also reported a 3x median speed gain, but METR flagged that self-reported estimates have previously overstated AI's effect by 40 percentage points against controlled measurement.
Benchmarks & progress · Culture & impact
METR: many SWE-bench-passing pull requests would not actually be merged
Four maintainers reviewing 296 AI-generated pull requests for scikit-learn, Sphinx and pytest found roughly half of automated-grader 'passes' would be rejected in real review.
Benchmarks & progress
METR updates time-horizon estimates (1.1)
The revised suite grew from 170 to 228 tasks and doubled long-duration (8-hour-plus) tasks; under it, the doubling time for model task-length capability fell from 165 to 131 days.
Benchmarks & progress
Anthropic issues a pilot sabotage risk report for Claude
Reviewed internally and by METR, the report found Claude Opus 4's risk of undetected sabotage 'very low, but not completely negligible.'
Safety & alignment
METR examines how time horizon varies across domains
Applying its 50%-success task-length method to nine benchmarks, METR found doubling times of two to six months for reasoning tasks but around twenty months for Tesla's self-driving system.
Benchmarks & progress
METR publishes 'Measuring AI Ability to Complete Long Software Tasks'
Introduced the 'time horizon' metric — task length a model can complete autonomously at 50% success — and found it doubling roughly every seven months.
Ideas & essays · Benchmarks & progress
METR reports frontier models show dangerous capability before public deployment
METR argued that model theft, internal misuse and misaligned agents pose risks during training and internal deployment, before any public release.
Safety & alignment
METR publishes a rogue AI replication threat-model report
Analysis finds no decisive technical barrier preventing a sufficiently capable model from self-replicating at scale outside lab control.
Safety & alignment
Also mentioned in 2 entries
Referenced in passing — METR isn't the main subject of these.