Benchmarks · Aggregate indices & arenas

LMArena

also: Chatbot Arena, LMSYS Chatbot Arena, Arena

Which of two anonymous models people prefer in open-ended, head-to-head conversation, aggregated into a running Elo rating rather than a fixed test score.

LMSYS / UC Berkeley, later independent as LMArena (Arena Intelligence)Released 3 May 2023Disputed

Most benchmarks score a model against a fixed answer key. LMArena, launched by LMSYS researchers at UC Berkeley as Chatbot Arena in May 2023, asks a different question: given two anonymous replies to the same prompt, which one does a person actually prefer? Visitors vote blind, the identities are revealed afterwards, and the votes accumulate into a running Elo-style rating rather than a one-off score — closer to how the models are used in practice, but harder to audit than a test with a right answer.

The approach made LMArena one of the most closely watched public rankings in the field, cited in launch announcements from every major lab as the votes piled into the millions. That prominence also made it a target. In April 2025 it emerged that Meta had submitted a specially tuned Llama 4 Maverick variant that outperformed the model it actually shipped, and weeks later “The Leaderboard Illusion” argued that undisclosed private pre-release testing structurally favoured well-resourced labs. LMArena disputed several of the paper’s figures but adopted disclosure and provisional-scoring policies in response.

Leadership has continued to change hands quickly since — Google’s Gemini 2.5 Pro and later Gemini 3 Pro both took the top spot on release, and by August 2026 the text arena, now rebranded simply “Arena,” had logged more than 7.7 million votes across 389 models, with Anthropic, Google and OpenAI systems clustered within a few Elo points at the top. It remains a widely cited signal of how models are received in open-ended use, but after 2025’s disputes it is generally read alongside fixed benchmarks rather than in place of them.

The set

Visitors submit a prompt to two randomly paired, anonymised models, read both replies and vote for the better one before either model's identity is revealed. Votes accumulate into a Bradley-Terry/Elo-style rating; as of August 2026 the text arena alone has logged over 7.7 million votes across 389 models.

Where it stands

Still one of the most-cited public rankings, but 2025's 'Leaderboard Illusion' paper and the Llama 4 Maverick episode left its scores widely treated as suggestive rather than authoritative.

How the top score changed hands

  1. May 2023Vicuna-13Bleader of the initial ~4,700-vote rankingThe launch leaderboard covered nine mostly open-weight, LLaMA-derived models; frontier proprietary models were not yet included.
  2. March 2025Gemini 2.5 Proled by what Google called 'a significant margin'Google's own release claim; came after roughly two years in which OpenAI and Anthropic models had generally led the board.
  3. November 2025Gemini 3 Pro1501 EloReported alongside record scores on GPQA Diamond and SWE-bench Verified at launch.

Current best: Claude Fable 5 — 1506 Elo Overall text leaderboard snapshot; rankings shift week to week and several models cluster within a few Elo points of the top.

In the timeline · 22 entries · showing 16 most notable

More aggregate indices & arenas benchmarks