The years

2024: reasoning, recognition, and real-world harm

Models learned to think for longer, AI's scientific work won two Nobel Prizes, and the EU's rulebook came into force — as deepfakes reached elections, chatbot-harm lawsuits began, and California's safety bill was vetoed.

Capability grew along new axes in 2024. GPT-4o brought fast, spoken, multimodal conversation; Anthropic’s Claude 3 and 3.5 and Google’s Gemini traded the lead through the year, with Gemini stretching to a million-token context window. The most consequential shift came in September, when OpenAI’s o1 showed that letting a model spend more time “thinking” at the moment of answering — rather than simply making it bigger — could sharply improve its reasoning, opening a second path to progress just as pre-training gains looked harder to come by. AI’s scientific standing was formally recognised: Geoffrey Hinton and John Hopfield shared the Nobel Prize in Physics for the foundations of neural networks, and Demis Hassabis and John Jumper shared the Chemistry prize for AlphaFold, whose third version now modelled interactions across proteins, DNA and drugs. The first genuine agents appeared, too, as Claude gained the ability to use a computer and Anthropic proposed a shared standard for connecting models to external tools.

The harms grew more concrete alongside the capabilities. Deepfakes moved from novelty to weapon: a fabricated robocall imitating President Biden told New Hampshire voters to stay home, and a video call with AI-generated colleagues was used to steal twenty-five million dollars from one company. In one of the year’s most sobering cases, a mother sued the chatbot service Character.AI after her teenage son’s death, beginning a wave of litigation over the effects of these systems on vulnerable users. On governance, the limits of ambition showed: California’s SB 1047, which would have imposed safety obligations on the largest models, was vetoed in September after heavy industry lobbying, and at OpenAI the dissolution of its “superalignment” team and a run of senior departures raised doubts about whether the leading labs would fund safety at the pace they funded scale. The EU AI Act entered into force, the first broad, binding regime — though its real effects lay in the years ahead.

The headlines of 2024

See every entry from 2024 on the timeline →

Benchmarks introduced in 2024

The other half of progress: as older tests saturate, new ones are built to stretch the frontier again.

Browse the full benchmark catalogue →