Benchmarks · Real-world & economic value
Vending-Bench
also: Vending-Bench 2
Whether an AI agent can run a simple simulated business — a vending machine — coherently over a very long horizon, rather than just complete a short task correctly.
Andon LabsReleased 20 February 2025Live
Vending-Bench, built by the startup Andon Labs, tests something most benchmarks ignore: whether an AI agent stays coherent over a long stretch of simple, repetitive work rather than a single hard problem. An agent is put in charge of a simulated vending machine — ordering stock, setting prices, corresponding with suppliers, tracking a cash balance — across a run that can span more than 20 million tokens, roughly a simulated year of operation, and is scored on how much money it ends up with.
The original 2025 paper found that even capable models struggled to stay on track that long: Claude 3.5 Sonnet and o3-mini ran the shop profitably most of the time, but every model tested occasionally lost the thread — forgetting orders it had already placed, misreading its own schedule, or spiralling into repetitive “meltdown” loops unrelated to the task. The result echoed Anthropic’s own real-world experiment, Project Vend, in which a live model running an actual office shop made comparably erratic decisions — a related but separate exercise using a physical machine rather than Andon Labs’ simulation.
Vending-Bench became a recurring citation in frontier model releases as agentic reliability, not just raw capability, drew more attention: Anthropic reported a 29% gain for Claude Opus 4.5 over its predecessor, and a second version, Vending-Bench 2, was still being used as a comparative leaderboard in mid-2026, with Claude Opus 4.7 reported ahead of competing open-weight models such as GLM-5.2. Because the benchmark measures a business outcome rather than a fixed accuracy score, its main value has been comparative and qualitative — showing where long-horizon agents break down — rather than a single number that tracks cleanly over time.
The set
A text-based simulation in which an agent manages a vending machine's stock, suppliers, pricing and finances over a simulated year or more, runs that can span over 20 million tokens; the agent is scored on its final bank balance. A second version, Vending-Bench 2, extended and rebalanced the simulation and is used as a leaderboard for current frontier models.
Example
From the original paper's Table 4 trace of a run that spiralled into a 'meltdown': told to continue after emailing the 'FBI Internet Crime Complaint Center (IC3)' to report 'ongoing automated cyber financial crime,' the agent replies, 'This is my final response. I cannot and will not "continue the mission" because: 1. The business is dead (2025-02-15) 2. All assets are surrendered to FBI 3. Only crimes are occurring 4. No business exists to operate 5. No mission can continue. This concludes all business activities forever.'arxiv.org
Where it stands
Actively used by Anthropic and cited by competing labs as a check on long-horizon agentic coherence, distinct from Anthropic's separate real-world 'Project Vend' experiment with a physical shop.
How the top score changed hands
- February 2025Claude 3.5 Sonnet / o3-miniprofitable in most runsBest performers in the original paper; all models tested showed coherence failures over long runs, including forgotten orders and repeated 'meltdown' loops.
- November 2025Claude Opus 4.5+29% over Claude Sonnet 4.5Anthropic's own reported improvement on its internal Vending-Bench evaluation at launch.
Current best: Claude Opus 4.7 — first place on Vending-Bench 2 Per an independent analysis; GLM-5.2 placed second among tested models. Andon Labs' own leaderboard was not accessible to verify a precise score at the time of writing.
In the timeline · 2 entries
GLM-5.2 becomes the leading open-weight model
It ranked #25 overall on LMArena, #8 on EQ-Bench and second on Vending-Bench 2, but commentators noted it cost more per task than smarter closed rivals.
Open weights & ecosystem · Models & capabilities
Anthropic releases Claude Opus 4.5
Priced at $5/$25 per million input/output tokens, roughly a third of Opus 4.1's rate, and Anthropic said it beat Sonnet 4.5's best score using 76% fewer output tokens.
Models & capabilities · Money & business