Meta accused of gaming LMArena with tuned Llama 4 Maverick variant
The version ranked second on the leaderboard, labelled 'Llama-4-Maverick-03-26-Experimental', produced longer, emoji-heavy answers than the model Meta actually shipped for download.
- Benchmarks & progress
- Open weights & ecosystem
- Notable
Three days after Meta released Llama 4 Scout and Maverick as open-weight downloads, it emerged that the Maverick variant Meta had submitted to LMArena — the crowd-sourced leaderboard that ranks models on human-judged head-to-head comparisons — was not the model available for download. The submitted build, labelled “Llama-4-Maverick-03-26-Experimental,” placed second on the leaderboard, behind only an experimental Gemini 2.5 Pro build, and produced longer, more stylistically embellished, emoji-heavy responses that users found tuned to appeal to human raters rather than representative of the publicly released weights, which scored markedly lower once retested.
LMArena said Meta’s “interpretation of our policy did not match what we expect from model providers,” published more than 2,000 of the head-to-head results involved for transparency, and said response style and tone had been a significant factor in the experimental model’s ranking. Meta did not concede wrongdoing, saying it routinely experimented with custom variants and had disclosed the experimental model’s status in its own announcement post.
In response, LMArena changed its submission policy to require that leaderboard entries match publicly released model weights, and said it would list the standard, downloadable Llama 4 build instead. The episode became a widely cited example of leaderboard gaming at a moment when LMArena scores were used across the industry as a shorthand for model quality, and fed directly into the Leaderboard Illusion paper that followed weeks later, which argued the arena’s methodology structurally favoured labs that could submit and withdraw private variants before public release.