Benchmarks · Aggregate indices & arenas

Open LLM Leaderboard

also: Hugging Face Open LLM Leaderboard, Open LLM Leaderboard v2

How an open-weight language model scores on a fixed suite of automated academic benchmarks, run and standardised by Hugging Face rather than self-reported by the model's developer.

Hugging FaceReleased Q2 2023Retired

For much of 2023 and 2024, a new open-weight model’s standing was measured against one place: Hugging Face’s Open LLM Leaderboard, a public ranking that ran submitted models through a fixed set of automated academic benchmarks via the EleutherAI evaluation harness and published the results without the developer able to cherry-pick which numbers to show. It existed by Falcon-40B’s release in May 2023, where the leaderboard’s top spot served as the third-party evidence behind the model’s “best open-source model” claim, and Hugging Face’s own team used it as a research platform, including a 2023 study documenting how GPT-4 behaved as an automated judge — finding it favoured longer answers and outputs resembling its own training data, a caution that shaped how model-graded evaluation was treated across the field afterward.

The leaderboard’s fixed benchmark suite was also its expiry date. As models improved, scores compressed toward the ceiling on the original tests, and in 2024 Hugging Face replaced them with a harder second-generation suite — MMLU-Pro, GPQA, MuSR, IFEval and BBH among them — to restore separation between models. Even that proved temporary: as reasoning-tuned, assistant-style models became the norm, Hugging Face concluded that static multiple-choice and knowledge tests no longer captured the behaviour that mattered to users.

Hugging Face retired the leaderboard in March 2025, after it had ranked more than 13,000 models across its two versions, rather than build a third static suite it expected to saturate again. It did not name a single successor, pointing instead to a fragmented landscape of more than 200 community-run, domain-specific leaderboards on its platform — an acknowledgment that no single fixed benchmark was likely to serve the open-model ecosystem the way it once had.

The set

The original version scored models on four evaluations run through the EleutherAI LM Evaluation Harness (Hugging Face's own June 2023 methodology post discusses these in depth). A 2024 second iteration replaced the suite with harder tests — including MMLU-Pro, GPQA, MuSR, IFEval and BBH — after the original set began to saturate. Over its life the leaderboard evaluated more than 13,000 open models across the two versions before being retired in March 2025.

Where it stands

Retired by Hugging Face in March 2025, which said fixed multiple-choice and knowledge tests could no longer meaningfully distinguish reasoning-era models; the company pointed users to a fragmented ecosystem of 200-plus community leaderboards instead of a single successor.

How the top score changed hands

  1. May 2023Falcon-40Btop of the leaderboard on releaseCited by developer TII's own model card as evidence of the claim that Falcon-40B was 'the best open-source model currently available.'

In the timeline · 6 entries

More aggregate indices & arenas benchmarks