Timeline

MMLU-Pro benchmark paper released

The paper reported chain-of-thought reasoning helped on the new benchmark where it had made little difference on the original MMLU, and cut prompt-sensitivity from 4-5 points to about 2.

  • Benchmarks & progress
  • Minor

Researchers at TIGER-AI-Lab published MMLU-Pro, a successor to the widely used MMLU benchmark, built to address the original test’s saturation as leading models approached or exceeded 90% accuracy. The new benchmark comprised more than 12,000 curated questions across 14 subject areas, drawn from academic exams and textbooks and filtered to remove trivial or noisy items from the original pool.

The central change was expanding each question’s answer choices from four to ten, sharply cutting the odds of guessing correctly, and shifting the question mix toward multi-step reasoning rather than recall. Against this harder test, the paper reported that model accuracy fell by 16 to 33 percentage points compared with the same models’ MMLU scores, and that chain-of-thought prompting produced a meaningful accuracy gain — unlike on the original MMLU, where direct answering had performed comparably. Sensitivity to how a question was phrased also fell, from a 4-5 point swing on MMLU to roughly 2 points on MMLU-Pro, suggesting the new benchmark measured something closer to underlying capability.

MMLU-Pro, accepted to NeurIPS 2024’s Datasets and Benchmarks track, became one of several successor benchmarks published in this period as MMLU itself lost its power to distinguish frontier models from each other — part of a recurring cycle in which a benchmark’s usefulness erodes as soon as it becomes a target for training and marketing.