OSWorld benchmarks AI agents on real desktop computer tasks
369 tasks across real Ubuntu, Windows and macOS applications, graded on the machine's actual end-state; the best model at release solved 12% against a 72% human baseline.
- Benchmarks & progress
- Notable
A team from the University of Hong Kong, Salesforce Research, Carnegie Mellon and the University of Waterloo released OSWorld, a benchmark that drops an AI agent into a whole real operating system — a working Ubuntu, Windows or macOS environment with genuine applications — and asks it to complete open-ended tasks by operating the graphical interface, via screenshots and clicks, rather than through a tidy API or a sandboxed browser. Its 369 tasks are graded by execution-based checker functions that inspect the machine’s actual end-state, so an agent is judged on whether the file was really reformatted or the setting really changed.
The difficulty was the point. At release, the strongest system solved about 12% of tasks against a human baseline near 72%, a gap that made OSWorld a standard yardstick for the computer-use agents that followed. When Anthropic gave Claude the ability to use a computer later in 2024 and OpenAI launched Operator in early 2025, both reported their OSWorld scores as the headline measure of progress.
Those scores rose quickly — frontier models were matching or beating the human baseline on the original tasks by 2026 — prompting the same group to release a harder OSWorld 2.0 with much longer tasks, on which even the strongest systems again fell back to around 20%. It became the clearest single illustration of how fast desktop-agent capability moved, and how quickly each version of the test was outgrown.