Benchmarks · Agents, tools & computer use
OSWorld
also: OSWorld-Verified, OSWorld 2.0
Can an agent operate a real desktop computer — a whole Ubuntu, Windows or macOS environment with real applications — to complete an open-ended task, rather than a sandboxed browser or single app?
University of Hong Kong, Salesforce Research, Carnegie Mellon University & University of WaterlooReleased 11 April 2024Live
OSWorld, released in April 2024 by researchers at the University of Hong Kong, Salesforce, Carnegie Mellon and Waterloo, tests something narrower and harder than most web-agent benchmarks: can a model operate an entire computer? Its 369 tasks run inside real Ubuntu, Windows and macOS environments with genuine applications — spreadsheets, file managers, browsers, settings panels — and are graded by checking the machine’s actual end-state afterwards, so an agent has to see the screen, decide what to click, and get the real result right.
At launch the best model managed only 12% against a 72% human baseline, and computer-use agents remained a research curiosity through most of 2024. That changed fast once labs began shipping dedicated computer-use products: OpenAI’s Operator reached 38% in January 2025, Claude Sonnet 4.5 hit 61% by September, and by early 2026 successive point releases from OpenAI and Google were reporting scores in the 65–83% range on the refreshed “OSWorld-Verified” task set — with OpenAI’s GPT-5.4 the first to claim it had passed the human baseline outright.
That pace of saturation prompted the same research group to release OSWorld 2.0 in June 2026: 108 much longer tasks, many taking a skilled human well over an hour, covering seven professional domains rather than isolated GUI actions. Early results reset the difficulty curve — even the strongest tested model cleared only about a fifth of tasks — reprising the pattern seen across the field’s other benchmarks, where a measure saturates and is promptly replaced by a harder one built the same way.
The set
369 tasks (361 excluding ones with an external Google Drive dependency) spanning real desktop and web applications across Ubuntu, Windows and macOS, graded by execution-based checker functions against the machine's actual end-state. A July 2025 refresh, OSWorld-Verified, fixed community-reported errors in the task set and added faster AWS-based parallel evaluation; the version most releases now cite.
Example
Rename "Sheet 1" to "LARS Resources". Then make a copy of it. Place the copy before "Sheet 2" and rename it by appending a suffix "(Backup)"arxiv.org
Where it stands
Original OSWorld tasks moved from a 12% ceiling in 2024 to frontier models matching or beating the ~72% human baseline by mid-2026, prompting the same group to release OSWorld 2.0 in June 2026 with 108 much longer tasks (median 1.6 hours of human time) — on which even the strongest models clear only around 20%.
How the top score changed hands
- April 2024Best model at release (original paper)12.24%Against a 72.36% human baseline on the same tasks.
- January 2025OpenAI Operator (CUA)38.1%Against 22.0% for Anthropic's Computer Use and a 72.4% human baseline.
- September 2025Claude Sonnet 4.561.4%
- February 2026GPT-5.3-Codex64.7%On OSWorld-Verified, roughly double GPT-5.2-Codex's 38.2%.
- February 2026Claude models (Anthropic's own figure)72.5%Cited by Anthropic when announcing its acquisition of Vercept, up from under 15% in late 2024.
- March 2026GPT-5.475%On OSWorld-Verified; OpenAI framed this as beating its own ~72.4% human baseline figure.
- July 2026Gemini 3.6 Flash83.0%
Current best: Gemini 3.6 Flash — 83.0% On OSWorld-Verified, as reported in Google's release announcement; against a roughly 72% human baseline on the same task set.
In the timeline · 11 entries
Google releases Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber
Google cut per-token cost and lifted coding and computer-use benchmark scores on its efficiency tier, while Gemini 3.5 Pro remained unreleased and still in partner testing.
Models & capabilities
Anthropic launches Claude Sonnet 5
Priced at $3/$15 per million input/output tokens against Opus 4.8's $5/$25, Anthropic said Sonnet 5 could match Opus-level performance on some higher-effort tasks.
Models & capabilities
OpenAI releases GPT-5.4
OpenAI's first general-purpose model with built-in computer-use, reported scoring 75% on OSWorld-Verified against 47.3% for GPT-5.2 and roughly 72% for human testers.
Models & capabilities
Anthropic acquires Vercept
Terms were undisclosed; Vercept will wind down its own product, and Anthropic cited Claude's OSWorld computer-use score rising from under 15% in late 2024 to 72.5%.
Money & business · Labs & people
Anthropic releases Claude Sonnet 4.6
Early testers preferred it to Sonnet 4.5 on coding tasks about 70% of the time, and to the larger Opus 4.5 about 59% of the time, at unchanged Sonnet pricing.
Models & capabilities
OpenAI releases GPT-5.3-Codex
OpenAI reported the model roughly doubled its predecessor's OSWorld-Verified computer-use score, from 38.2% to 64.7%, and was the first Codex model rated 'High capability' for cybersecurity tasks.
Models & capabilities
Anthropic ships Claude Sonnet 4.5
Anthropic reported 77.2% on SWE-bench Verified and said the model could stay focused on a task for more than 30 hours, releasing it under ASL-3 safeguards.
Models & capabilities
METR examines how time horizon varies across domains
Applying its 50%-success task-length method to nine benchmarks, METR found doubling times of two to six months for reasoning tasks but around twenty months for Tesla's self-driving system.
Benchmarks & progress
OpenAI launches Operator
Built on a new Computer-Using Agent model layered on GPT-4o, it scored 38.1% on OSWorld against a 72.4% human baseline, and launched to $200-a-month Pro subscribers only.
Models & capabilities
Claude gets computer use
The public beta let Claude view screenshots and issue cursor, click and keystroke commands, scoring 14.9% on OSWorld against 7.8% for the nearest rival.
Models & capabilities
OSWorld benchmarks AI agents on real desktop computer tasks
369 tasks across real Ubuntu, Windows and macOS applications, graded on the machine's actual end-state; the best model at release solved 12% against a 72% human baseline.
Benchmarks & progress