Benchmarks · Aggregate indices & arenas
LiveBench
How a model performs on recently created questions, graded by objective ground-truth answers rather than human or LLM judgment, so that scores cannot reflect memorised test data and cannot be inflated by a biased judge model.
Abacus.AI, NYU, University of Maryland and collaborators (Colin White, Samuel Dooley, Tom Goldstein, Yann LeCun and others)Released 24 June 2024Live
LiveBench was built to answer a specific worry about every other benchmark on this list: that a widely used test eventually leaks into a model’s training data, so a high score reflects memorisation rather than ability. A team spanning Abacus.AI, NYU and the University of Maryland — including Tom Goldstein and Yann LeCun among its authors — published it in June 2024 with a fix: draw questions from sources too recent to have been trained on, such as that month’s arXiv papers, news articles and maths competitions, refresh the set on a rolling schedule, and grade every answer against an objective ground truth rather than a human or LLM judge’s opinion.
That design also sidesteps a second problem, the one that dogs arenas and LLM-judged benchmarks alike: a judge model can be gamed by style rather than substance. LiveBench’s 23 tasks, spread across reasoning, coding, agentic coding, mathematics, data analysis, language and instruction following, are checked against fixed answers instead. The tradeoff is that the whole question set is periodically replaced — the site had moved from its original 2024-06-24 release to a 2026-06-25 edition by the time of writing — which keeps scores comparison-resistant across versions in the way a static benchmark’s are not.
On the current release, Anthropic’s Claude Fable 5 led with an overall score of 83.0, narrowly ahead of OpenAI’s GPT-5.6 Sol at 81.0, with cost-per-successful-task published alongside every score so a reader can weigh capability against price directly. Because each refresh changes the underlying questions, LiveBench functions less as a fixed yardstick than as a standing methodology for building one — a response to benchmark contamination rather than a single benchmark in the usual sense.
The set
23 objective tasks across seven categories — reasoning, coding, agentic coding, mathematics, data analysis, language and instruction following — drawn from sources such as recent math competitions, arXiv papers and news articles. Every task is scored against a ground-truth answer rather than by human or LLM preference. The full question set is refreshed roughly every six months (the latest as of August 2026 dated 2026-06-25) specifically to stay ahead of training-data contamination, and older releases remain browsable for comparison.
Where it stands
Actively maintained with periodic full refreshes rather than one static set; per-task cost is published alongside scores.
In the timeline · 2 entries
Alibaba releases Qwen2.5-Max
Unlike most of Alibaba's Qwen line, Max was released as a proprietary API-only model, pretrained on over 20 trillion tokens, which Alibaba said beat DeepSeek-V3 on several benchmarks.
Models & capabilities · Benchmarks & progress
StepFun launches Step-2, a trillion-parameter MoE model
StepFun said the mixture-of-experts model approximated GPT-4 on maths, logic, coding and dialogue; it was unveiled alongside a multimodal and an image-generation model at WAIC.
Models & capabilities · Open weights & ecosystem