
Organisation
LMArena
Public leaderboard, originally the Berkeley-linked Chatbot Arena, that ranks AI models by anonymous head-to-head human votes rather than fixed test sets.
LMArena, which began as Chatbot Arena from the Berkeley-linked LMSYS group in 2023, is a public leaderboard that ranks AI models by anonymous, head-to-head human votes rather than fixed test sets. The format — people picking the better of two randomly paired responses, aggregated into an Elo-style rating — became one of the field's most closely watched rankings, cited by labs in launch announcements and by journalists as an informal measure of the state of the art. That influence made it a target as much as a reference: in 2025 the "Leaderboard Illusion" paper argued that undisclosed private testing let well-resourced labs quietly tune models against the leaderboard, and Meta was separately accused of submitting a model that differed from the one it shipped. LMArena disputed several of the specific claims while adopting new disclosure policies, and the dispute fed a wider shift toward benchmarks that are harder to game.
- Category
- Evaluation & forecasting
- Founded
- 2023
- HQ
- San Francisco, US
- Funding
- $100m seed led by a16z and UC Investments (2025)
- Key people
- Anastasios Angelopoulos, Wei-Lin Chiang, Ion Stoica
Appears alongside
Featured in threads
Tracks
- Benchmarks & progress 5
- Open weights & ecosystem 2
- Models & capabilities 1
LMArena responds to 'Leaderboard Illusion' paper with policy changes
LMArena disputed the paper's headline figures on open-model share and score-boosting but agreed to mark scores 'provisional' and disclose pre-release testing.
Benchmarks & progress
'The Leaderboard Illusion' paper critiques Chatbot Arena methodology
Researchers found Meta tested roughly 27 private Llama variants before its public release and that OpenAI and Google alone received about 40% of all Arena battle data.
Benchmarks & progress
Meta accused of gaming LMArena with tuned Llama 4 Maverick variant
The version ranked second on the leaderboard, labelled 'Llama-4-Maverick-03-26-Experimental', produced longer, emoji-heavy answers than the model Meta actually shipped for download.
Benchmarks & progress · Open weights & ecosystem
Llama 4 lands badly
Meta's mixture-of-experts release was undercut by accusations that a version tuned for LMArena differed from the public weights.
Open weights & ecosystem · Benchmarks & progress · Models & capabilities
LMSYS launches Chatbot Arena
The Berkeley-linked project ranked chatbots by anonymous, randomised head-to-head votes rather than fixed test sets, and later became LMArena.
Benchmarks & progress
Also mentioned in 13 entries
Referenced in passing — LMArena isn't the main subject of these.
- June 2026GLM-5.2 becomes the leading open-weight model
- January 2026Baidu launches ERNIE 5.0, a 2.4-trillion-parameter native multimodal model
- December 2025Mistral launches Mistral 3 model family
- November 2025Google ships Gemini 3
- November 2025xAI releases Grok 4.1
- June 2025Google updates Gemini 2.5 Pro preview with improved coding performance
- March 2025Gemini 2.5 Pro takes the lead on reasoning benchmarks
- March 2025Google releases Gemma 3, an open model family built on Gemini 2.0
- February 2025xAI releases Grok-3
- December 2024Google releases Gemini 2.0 Flash Thinking, its first public reasoning model
- November 2024AI2 releases Tulu 3 post-training recipe
- August 2024xAI releases Grok-2
- June 2023vLLM releases PagedAttention inference engine