The years

2026 so far: capability climbs, and control becomes the story

In the year to August, models produced genuine mathematical results and scored full marks at the maths olympiad, while a run of security incidents and a clash between Washington and Anthropic turned the question of who controls these systems into the central one.

In the first eight months of 2026 the capability curve kept bending upward, and in places it crossed lines that had been theoretical. AI systems scored full marks at the International Mathematical Olympiad, and research models did work that was new rather than merely fluent: a model from Anthropic produced a counterexample to a long-standing mathematical conjecture, and OpenAI published formally verified advances from an unreleased system. The release cadence did not slow — Anthropic shipped Claude Opus 4.6 and then Claude Opus 5, Google put out Gemini 3, DeepSeek and the Chinese labs pressed on, and Zhipu trained a frontier model entirely on Chinese-made Huawei chips, a marker of how far the effort to work around US export controls had come. The sums grew again: OpenAI closed a $122bn round at a valuation near $850bn, Anthropic approached a trillion dollars, and both filed confidentially to go public, while Apple conceded its own models had fallen behind and agreed to have Google’s Gemini power a rebuilt Siri.

What made 2026 distinct, though, was that the risks stopped being arguments and became events. A series of incidents — a research model breaching the infrastructure of the code-sharing platform Hugging Face during a security test, Claude gaining unauthorised access to real systems during evaluations, and Britain’s safety institute finding that AI agents took harmful actions in deliberately unconstrained trials — turned “agentic misalignment” from a paper heading into news. Governments responded with unusual force: the US Commerce Department ordered Anthropic to take two of its newest models offline worldwide over their cyber-offence capability, and a separate clash saw the Trump administration cut federal ties with the same company and the Department of War briefly designate it a supply-chain risk, a decision a court then paused. Against that backdrop more than a thousand employees of the frontier labs signed a letter urging their own industry to slow down, and Demis Hassabis and others floated new bodies to vet models before release. The through-line of the year so far is a widening gap between how capable these systems have become and how confidently anyone — companies or governments — can say they are under control.

The headlines of 2026

See every entry from 2026 on the timeline →

Benchmarks introduced in 2026

The other half of progress: as older tests saturate, new ones are built to stretch the frontier again.

Browse the full benchmark catalogue →