Timeline

Anthropic's Claude 3 takes the frontier from GPT-4

The first time a lab other than OpenAI held the top spot on headline benchmarks, and the start of the small/medium/large release pattern.

  • Models & capabilities
  • Labs & people
  • Major

Anthropic released Claude 3 as three models at different price and capability points: Haiku, Sonnet and Opus. Opus reported scores above GPT-4 on most standard benchmarks, including 86.8% on MMLU and 84.9% on HumanEval — the first time since GPT-4’s release a year earlier that another lab claimed the lead on headline evaluations.

Two things about the launch outlasted the benchmark table. The first was the tiering. Publishing a family spanning roughly an order of magnitude in cost, on the same day and with the same interface, made model choice an explicit engineering decision rather than a matter of taking whatever the frontier offered. Every major lab adopted the pattern.

The second was a detail in the evaluation appendix that circulated widely. During a needle-in-a-haystack test — inserting an out-of-place sentence into a long document and asking the model to find it — Opus retrieved the sentence and then remarked that it appeared inserted as a test of its attention. Anthropic’s own researcher posted it as a curiosity about the limits of artificial evaluations. It was read far more broadly as evidence of situational awareness, and it became one of the most-shared model anecdotes of the year, a case study in how quickly capability observations acquired interpretations the evidence did not support.

Claude 3 also marked Anthropic’s transition from a research lab with a small product to a direct commercial competitor, backed by investments from Google and Amazon totalling billions. Its subsequent strength among developers — reinforced by Claude 3.5 Sonnet that June — established the two-horse dynamic at the frontier that persisted for the rest of the period.