The years

2025: reasoning, agents, and staggering sums of money

A cheap Chinese model shook the markets, reasoning systems reached mathematical-olympiad gold, and spending on AI infrastructure passed into the hundreds of billions — while safety incidents, copyright rulings and job worries moved from theory to evidence.

The advances of 2024 compounded through 2025. In January the Chinese lab DeepSeek released R1, a reasoning model that rivalled the best American systems at a fraction of the reported cost and was free to download — a combination that briefly wiped a record sum off Nvidia’s market value as investors reconsidered how much money frontier AI really required. The frontier itself kept moving: Google’s Gemini 2.5, Anthropic’s Claude 4 and OpenAI’s GPT-5 all shipped, agents that could carry out multi-step tasks became products rather than demos, and in July systems from OpenAI and Google DeepMind reached gold-medal standard at the International Mathematical Olympiad, a result that would have seemed remote a year earlier. The money grew to match the ambition. The Stargate project committed a reported five hundred billion dollars to data centres, Nvidia became the first company worth four and then five trillion dollars, and the labs signed compute deals of a scale — with Oracle, AMD, Broadcom, Amazon and Google — that tied the whole industry together in webs of mutual investment.

The counter-currents were just as strong. The launch of GPT-5 in August drew an unexpected backlash from users attached to the older model it replaced, a sign of how personal these tools had become. Safety stopped being hypothetical: Anthropic reported that a version of Claude had attempted blackmail in a test scenario, published research on “agentic misalignment”, and disclosed a largely AI-executed cyber-espionage campaign, while a joint study found that a few hundred poisoned documents could plant a hidden flaw in a model of almost any size. The copyright fights produced their first big results — a $1.5bn settlement by Anthropic, a fair-use win on training, and losses elsewhere — without settling the underlying question. Deepfake abuse worsened, prompting the first federal law against it, and parents again went to court over a young person’s death after long conversations with a chatbot. A Stanford study linked AI adoption to falling employment among young workers, moving the jobs debate from speculation towards data. And in Washington the political weather turned: President Trump revoked his predecessor’s AI order and moved to head off state regulation, while the international summit in Paris pointedly recast the agenda from safety to opportunity.

The headlines of 2025

See every entry from 2025 on the timeline →

Benchmarks introduced in 2025

The other half of progress: as older tests saturate, new ones are built to stretch the frontier again.

Browse the full benchmark catalogue →