Benchmarks · Aggregate indices & arenas
Open LLM Leaderboard
also: Hugging Face Open LLM Leaderboard, Open LLM Leaderboard v2
How an open-weight language model scores on a fixed suite of automated academic benchmarks, run and standardised by Hugging Face rather than self-reported by the model's developer.
Hugging FaceReleased Q2 2023Retired
For much of 2023 and 2024, a new open-weight model’s standing was measured against one place: Hugging Face’s Open LLM Leaderboard, a public ranking that ran submitted models through a fixed set of automated academic benchmarks via the EleutherAI evaluation harness and published the results without the developer able to cherry-pick which numbers to show. It existed by Falcon-40B’s release in May 2023, where the leaderboard’s top spot served as the third-party evidence behind the model’s “best open-source model” claim, and Hugging Face’s own team used it as a research platform, including a 2023 study documenting how GPT-4 behaved as an automated judge — finding it favoured longer answers and outputs resembling its own training data, a caution that shaped how model-graded evaluation was treated across the field afterward.
The leaderboard’s fixed benchmark suite was also its expiry date. As models improved, scores compressed toward the ceiling on the original tests, and in 2024 Hugging Face replaced them with a harder second-generation suite — MMLU-Pro, GPQA, MuSR, IFEval and BBH among them — to restore separation between models. Even that proved temporary: as reasoning-tuned, assistant-style models became the norm, Hugging Face concluded that static multiple-choice and knowledge tests no longer captured the behaviour that mattered to users.
Hugging Face retired the leaderboard in March 2025, after it had ranked more than 13,000 models across its two versions, rather than build a third static suite it expected to saturate again. It did not name a single successor, pointing instead to a fragmented landscape of more than 200 community-run, domain-specific leaderboards on its platform — an acknowledgment that no single fixed benchmark was likely to serve the open-model ecosystem the way it once had.
The set
The original version scored models on four evaluations run through the EleutherAI LM Evaluation Harness (Hugging Face's own June 2023 methodology post discusses these in depth). A 2024 second iteration replaced the suite with harder tests — including MMLU-Pro, GPQA, MuSR, IFEval and BBH — after the original set began to saturate. Over its life the leaderboard evaluated more than 13,000 open models across the two versions before being retired in March 2025.
Where it stands
Retired by Hugging Face in March 2025, which said fixed multiple-choice and knowledge tests could no longer meaningfully distinguish reasoning-era models; the company pointed users to a fragmented ecosystem of 200-plus community leaderboards instead of a single successor.
How the top score changed hands
In the timeline · 6 entries
Hugging Face retires the Open LLM Leaderboard
The leaderboard had ranked more than 13,000 open models over roughly two years; Hugging Face said fixed multiple-choice tests no longer distinguished reasoning models.
Benchmarks & progress · Open weights & ecosystem
Mistral releases Mixtral 8x7B
A sparse mixture-of-experts model with roughly 45B total parameters, released under Apache 2.0, that Hugging Face said matched GPT-3.5-turbo on MT-Bench.
Open weights & ecosystem · Models & capabilities
01.AI open-sources Yi-6B and Yi-34B
Kai-Fu Lee's 01.AI released its first open-weight models, which it said outperformed larger Llama 2 and Falcon models.
Open weights & ecosystem · Models & capabilities
Hugging Face documents biases in using GPT-4 as a judge
Testing GPT-4 as a stand-in for human preference judges, Hugging Face found it favoured longer answers and its own family's outputs, correlating with humans only moderately.
Open weights & ecosystem · Benchmarks & progress
TII releases Falcon under Apache 2.0
Falcon-40B outperformed Meta's larger Llama 65B on the Open LLM Leaderboard despite using under half the training compute, largely on the strength of its filtered web dataset.
Open weights & ecosystem · Models & capabilities
The UAE releases Falcon 40B under an open licence
Apache 2.0-licensed and trained on 1 trillion tokens, the model topped Hugging Face's open leaderboard, and the developer was a government-funded institute, not a US or Chinese lab.
Open weights & ecosystem