MMMU
also: Massive Multi-discipline Multimodal Understanding, MMMU-Pro
Can a model answer college-exam-level questions that genuinely require reading an accompanying image — a chart, diagram, map or chemical structure — rather than knowledge alone?
Ohio State University & University of Waterloo (Yue, Su, Chen et al.)Released 27 November 2023Live
MMMU asks whether a model can use an image, not just describe it. Each question comes from a real college exam, quiz or textbook and is genuinely unanswerable without reading the accompanying figure — a circuit diagram, a chemical structure, a music score, a map. Researchers at Ohio State University and the University of Waterloo built it that way deliberately, spreading 11.5K questions across six broad disciplines and 183 subfields so that no single visual skill could be gamed in isolation.
At launch in November 2023, the best available systems were well short of expert level: GPT-4V and Gemini Ultra scored 56% and 59% against a much higher human-expert baseline. Progress from there was rapid but shallow, and a year later the same team released MMMU-Pro to close the gap: it filters out any question a text-only model could already answer, expands the multiple-choice options, and adds a “vision-only” mode that renders the whole question as a single image, closing off the option of reading past the picture entirely. Reported scores dropped by 17 to 27 percentage points across the same models.
MMMU-Pro has since become one of the standard multimodal citations in frontier model announcements. Scores that sat around 51% for GPT-4o, Claude 3.5 Sonnet and Qwen2.5-VL-72B in February 2025 had climbed to the low 80s by the end of the year, with Gemini 3 Pro reporting 81.0% and the cheaper Gemini 3 Flash edging it out at 81.2% a few weeks later. Unlike some earlier multimodal benchmarks, it has not obviously saturated, and continues to sit alongside newer, harder successors such as CharXiv Reasoning in the eval suites frontier labs report at release.
The set
11.5K multimodal questions drawn from college exams, quizzes and textbooks, spanning six disciplines (Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, Tech & Engineering), 30 subjects and 183 subfields, with 30 heterogeneous image types. A 2024 follow-up, MMMU-Pro, tightens the test: it filters out questions a text-only model can already answer, expands the multiple-choice options, and adds a vision-only setting where the whole question is embedded as one image, closing off OCR shortcuts.
Example
Consider the following balance sheet for TD (with an accompanying image). Suppose that someone deposited $700 at TD Bank. Given this data, what is the minimum amount by which the money supply will increase? A) 0 B) 700 C) 1400 D) 3418 (Answer: A — 0)huggingface.co
Where it stands
MMMU-Pro remains part of frontier labs' own reported eval suites as of late 2025 (Gemini 3 Pro, Gemini 3 Flash); scores have climbed from the 50s in early 2025 to the low 80s by the end of the year, so it has not yet saturated the way the original MMMU largely has.
How the top score changed hands
- November 2023GPT-4V / Gemini Ultra56% / 59% (MMMU)The best-performing systems at launch, on a benchmark designed so that human experts still substantially outscore them.
- February 2025GPT-4o / Claude 3.5 Sonnet / Qwen2.5-VL-72B51.9% / 51.5% / 51.1% (MMMU-Pro)MMMU-Pro's harder vision-only and expanded-option format roughly halves scores compared with the original MMMU, per Qwen2.5-VL's own technical report.
- November 2025Gemini 3 Pro81.0% (MMMU-Pro)Google's own evaluation report, averaged across MMMU-Pro's Standard and Vision settings; GPT-5.1 scored 76.0% and Claude Sonnet 4.5 scored 68.0% on the same table.
Current best: Gemini 3 Flash — 81.2% (MMMU-Pro) From Google's official release announcement; edges out Gemini 3 Pro's self-reported 81.0% from three weeks earlier.
In the timeline · 4 entries
Google releases Gemini 3.1 Flash-Lite
Priced at $0.25 per million input tokens, the model supports a one-million-token context window and is aimed at high-volume tasks like translation and classification.
Models & capabilities
Moonshot AI releases Kimi K2.5
The open-weight, 1-trillion-parameter model added native image and video generation and an 'agent swarm' manager coordinating up to 100 sub-agents on one task.
Open weights & ecosystem · Models & capabilities
Google makes Gemini 3 Flash the default model across its products
Priced at $0.50/$3.00 per million tokens, Google reported it ran three times faster than Gemini 2.5 Pro while scoring 33.7% on Humanity's Last Exam, against 37.5% for Gemini 3 Pro.
Models & capabilities
Stanford HAI releases 2025 AI Index Report
The eighth annual report put US private AI investment at $109.1 billion in 2024, nearly twelve times China's $9.3 billion, and inference cost for GPT-3.5-level performance down over 280-fold since late 2022.
Benchmarks & progress · Ideas & essays