MathVista
Can a model do mathematical reasoning when the problem is given as a picture — a plot, a geometry diagram, a puzzle figure — rather than as text?
UCLA, University of Washington & Microsoft Research (Lu, Bansal, Galley, Gao et al.)Released 3 October 2023
Word problems are easy to make harder by taking away the words. MathVista, built by researchers at UCLA, the University of Washington and Microsoft Research, tests mathematical reasoning that depends on reading a picture first — a plotted function, a geometry diagram, a bar chart, a logic-puzzle figure — and only then doing the arithmetic or algebra the picture implies. The benchmark pools 6,141 examples from 28 existing maths-and-vision datasets plus three tasks the authors built themselves to cover gaps: puzzle figures, function plots and academic-paper diagrams.
At launch in October 2023, the gap between models and people was stark. GPT-4V, the best of twelve systems tested, scored 49.9%, comfortably ahead of Google’s Bard but still 10.4 points short of the 60.3% human baseline the authors measured. That gap closed fast: Claude 3.5 Sonnet reported 67.7% in June 2024, already above the original human baseline, and by February 2025 Alibaba’s Qwen2.5-VL-72B reported 74.8% on the same “testmini” subset, ahead of GPT-4o and Claude 3.5 Sonnet on the same comparison table in its own technical report.
Unlike some contemporaries, MathVista has not been superseded by a discrete “Pro” successor, but it has faded somewhat from frontier labs’ own headline evaluation reports; Google’s Gemini 3 Pro evaluation document from November 2025, for instance, reports a different, newer slate of multimodal and reasoning benchmarks and does not include a MathVista figure. Whether a specific 2025–26 frontier model has since posted a higher, independently sourced MathVista score than Qwen2.5-VL-72B’s 74.8% was not confirmed for this entry.
The set
6,141 examples combining 28 existing maths-related multimodal datasets with three tasks the authors built specifically for the benchmark: IQTest (logical reasoning over puzzle figures), FunctionQA (algebraic reasoning over plotted functions) and PaperQA (reasoning over figures from academic papers). Answers are multiple-choice or free-form numeric/text, checked by exact or normalised match; a 5,000-example 'testmini' subset is the one most models report.
Example
Is the function (f: R to R) injective? Choices: (A) Yes (B) No (Correct output: (B) No)arxiv.org
Where it stands
GPT-4V scored 49.9% at launch against a 60.3% human baseline; by February 2025 several models were reporting above that human baseline, with Qwen2.5-VL-72B at 74.8% on testmini — the most recent figure independently confirmed here, though later 2025-26 frontier releases plausibly score higher.
How the top score changed hands
- October 2023GPT-4V49.9% (testmini)Best of 12 models evaluated at launch, 15.1 points ahead of Google's Bard, but still 10.4 points short of the 60.3% human baseline.
- June 2024Claude 3.5 Sonnet67.7% (testmini)From Anthropic's own model card addendum, ahead of GPT-4o (63.8%) and Gemini 1.5 Pro (63.9%) on the same table.
- February 2025Qwen2.5-VL-72B74.8% (testmini)Above GPT-4o (63.8%) and Claude 3.5 Sonnet (67.7%) on the same comparison table.
Current best: Qwen2.5-VL-72B — 74.8% (testmini) Most recent score independently confirmed from an official technical report; several 2025-26 frontier labs have stopped headlining MathVista in release announcements, so a more recent leader may exist but isn't confirmed here.
In the timeline · 2 entries
Moonshot AI releases Kimi K1.5 reasoning model
Moonshot said its RL-trained model matched OpenAI's o1 on multimodal reasoning without Monte Carlo tree search, but it launched the same week as DeepSeek-R1 and drew far less attention.
Models & capabilities · Benchmarks & progress
xAI releases Grok-2
The beta release added image generation via Black Forest Labs' FLUX.1 and, within days, took second place on the LMSYS Chatbot Arena leaderboard behind GPT-4o.
Models & capabilities