Benchmarks · Agents, tools & computer use

OSWorld

also: OSWorld-Verified, OSWorld 2.0

Can an agent operate a real desktop computer — a whole Ubuntu, Windows or macOS environment with real applications — to complete an open-ended task, rather than a sandboxed browser or single app?

University of Hong Kong, Salesforce Research, Carnegie Mellon University & University of WaterlooReleased 11 April 2024Live

OSWorld, released in April 2024 by researchers at the University of Hong Kong, Salesforce, Carnegie Mellon and Waterloo, tests something narrower and harder than most web-agent benchmarks: can a model operate an entire computer? Its 369 tasks run inside real Ubuntu, Windows and macOS environments with genuine applications — spreadsheets, file managers, browsers, settings panels — and are graded by checking the machine’s actual end-state afterwards, so an agent has to see the screen, decide what to click, and get the real result right.

At launch the best model managed only 12% against a 72% human baseline, and computer-use agents remained a research curiosity through most of 2024. That changed fast once labs began shipping dedicated computer-use products: OpenAI’s Operator reached 38% in January 2025, Claude Sonnet 4.5 hit 61% by September, and through early 2026 successive releases pushed into the 65–75% range on the refreshed “OSWorld-Verified” task set — with OpenAI’s GPT-5.4 the first to claim it had passed the human baseline outright. By mid-2026 the top figures reached ~85%: Anthropic reported Claude Fable 5 and Mythos 5 at 85.0% in June, and later releases did not surpass it — Moonshot’s Kimi K3 reached 84.8% and Google’s Gemini 3.6 Flash 83.0%. Every one of these is a vendor’s own figure, since OSWorld-Verified has no single independent leaderboard.

That pace of saturation prompted the same research group to release OSWorld 2.0 in June 2026: 108 much longer tasks, many taking a skilled human well over an hour, covering seven professional domains rather than isolated GUI actions. Early results reset the difficulty curve — even the strongest tested model cleared only about a fifth of tasks — reprising the pattern seen across the field’s other benchmarks, where a measure saturates and is promptly replaced by a harder one built the same way.

The set

369 tasks (361 excluding ones with an external Google Drive dependency) spanning real desktop and web applications across Ubuntu, Windows and macOS, graded by execution-based checker functions against the machine's actual end-state. A July 2025 refresh, OSWorld-Verified, fixed community-reported errors in the task set and added faster AWS-based parallel evaluation; the version most releases now cite.

Example

Rename "Sheet 1" to "LARS Resources". Then make a copy of it. Place the copy before "Sheet 2" and rename it by appending a suffix "(Backup)"arxiv.org

Where it stands

Original OSWorld tasks moved from a 12% ceiling in 2024 to the top OSWorld-Verified figures reaching ~85% by mid-2026 (Claude Fable 5 / Mythos 5), well past the ~72% human baseline — prompting the same group to release OSWorld 2.0 in June 2026 with 108 much longer tasks (median 1.6 hours of human time), on which even the strongest models clear only around 20%. Every headline figure is a vendor self-report; there is no single independent OSWorld-Verified leaderboard.

Editions, and how each was led

OSWorld-2.1current

Released September 2026The 10 September 2026 update of the 2.0 task files, assets and companion web apps, which the official leaderboard lists as a separate release. Anthropic's runs on it also use a harness that keeps every screenshot and compacts context past 100k tokens; it says the results supersede its earlier 2.0 figures and are not comparable with them. So far every figure here is Anthropic's, including its re-runs of older models.

Best reported score at each point the lead changed hands. Tap a marker for the entry.

  1. June 2026Claude Sonnet 525.6% (strict)Anthropic's re-run in the Sonnet 5.5 system card.
  2. July 2026Claude Opus 537.2% (strict)Anthropic's re-run in the Opus 5.5 system card.
  3. September 2026Claude Fable 5.142.8% (strict)Anthropic's re-run in the Opus 5.5 system card (41.7% on the earlier task files).
  4. September 2026Claude Opus 5.548.7% (strict)The current top; Sonnet 5.5 followed at 43.5%.

Current best: Claude Opus 5.5 — 48.7% (strict) Opus 5.5 system card, max effort, five runs (81.8% partial). Under the same configuration: Sonnet 5.5 43.5%, Fable 5.1 42.8%, Opus 5 37.2%, Sonnet 5 25.6%.

OSWorld-2.1 (partial)

Released September 2026The 2.1 task files scored with partial credit — roughly 35–40 points above the strict score for the same model. Anthropic-only so far.

Best reported score at each point the lead changed hands. Tap a marker for the entry.

  1. June 2026Claude Sonnet 557.0% (partial)Anthropic's re-run in the Sonnet 5.5 system card.
  2. July 2026Claude Opus 574.0% (partial)Anthropic's re-run in the Opus 5.5 system card.
  3. September 2026Claude Fable 5.180.7% (partial)Anthropic's re-run in the Opus 5.5 system card (77.9% on the earlier task files).
  4. September 2026Claude Opus 5.581.8% (partial)The current top, just ahead of Fable 5.1 and Sonnet 5.5.

Current best: Claude Opus 5.5 — 81.8% (partial) Opus 5.5 system card and launch table, max effort (48.7% strict). Under the same configuration: Fable 5.1 80.7%, Sonnet 5.5 80.1%, Opus 5 74.0%, Sonnet 5 57.0%.

OSWorld-2.0

Released June 2026108 much longer tasks (median ~1.6 hours of human work) across seven professional domains, released June 2026 after Verified saturated. Under strict scoring even the strongest models clear only around 40%. (Vendors also report a partial-credit number roughly 35 points higher — a different metric, kept off this axis.)

Best reported score at each point the lead changed hands. Tap a marker for the entry.

  1. June 2026Claude Fable 536.1% (strict)An early point on the harder 2.0 set; Anthropic-reported, strict scoring.
  2. July 2026Claude Opus 539.6% (strict)Strict scoring, as listed in Anthropic's Fable 5.1 comparison table.
  3. September 2026Claude Fable 5.141.7% (strict)Strict scoring; 77.9% with partial credit. The top on the pre-September task files.

Current best: Claude Fable 5.1 — 41.7% (strict) Anthropic's Fable 5.1 launch table, strict scoring (77.9% with partial credit). Opus 5.5's 48.7%, briefly shown here, was run on the September task files with a changed harness and now sits on the 2.1 edition.

OSWorld-2.0 (partial)

Released June 2026The same 2.0 task set scored with partial credit rather than strict pass/fail — the number vendors more often headline, and now reported across OpenAI, Anthropic and Google, so it earns its own axis. Runs roughly 30–35 points above the strict score for the same model.

Best reported score at each point the lead changed hands. Tap a marker for the entry.

  1. June 2026Claude Fable 572.9% (partial)An early point on the 2.0 set, partial-credit; Anthropic-reported.
  2. September 2026Claude Fable 5.177.9% (partial)Partial-credit scoring; Anthropic-reported. The top on the pre-September task files.
  3. September 2026GPT-6 Astra72.6% (partial)OpenAI's launch table (offline set, v2026.08.08), partial score.

Current best: Claude Fable 5.1 — 77.9% (partial) Anthropic's Fable 5.1 launch table, partial-credit scoring (41.7% strict). Others on this edition: Opus 5 75.4%, GPT-6 Astra 72.6% and GPT-6.1 Sol 71.4% (OpenAI, v2026.08.08 offline set), GPT-5.6 Sol 65.7%, Gemini 3.8 Flash 59.0%. Opus 5.5's 81.8% was run on the September task files and moved to the 2.1 edition.

OSWorld-VerifiedSaturated

Released July 2025The 2024 original, refreshed in July 2025 to fix task errors and add faster parallel evaluation — the set most 2025–26 releases cite. Saturated past the ~72% human baseline by mid-2026, which prompted OSWorld 2.0.

Best reported score at each point the lead changed hands. Tap a marker for the entry.

  1. April 2024Best model at release (original paper)12.24%Against a 72.36% human baseline on the same tasks.
  2. January 2025OpenAI Operator (CUA)38.1%Against 22.0% for Anthropic's Computer Use and a 72.4% human baseline.
  3. September 2025Claude Sonnet 4.561.4%
  4. February 2026GPT-5.3-Codex64.7%On OSWorld-Verified, roughly double GPT-5.2-Codex's 38.2%.
  5. February 2026Claude models (Anthropic's own figure)72.5%Cited by Anthropic when announcing its acquisition of Vercept, up from under 15% in late 2024.
  6. March 2026GPT-5.475%On OSWorld-Verified; OpenAI framed this as beating its own ~72.4% human baseline figure.
  7. June 2026Claude Fable 5 / Mythos 585.0%OSWorld-Verified, Anthropic-reported — the top figure on this edition. Kimi K3 reached 84.8% and Gemini 3.6 Flash 83.0% a month or more afterwards.

Current best: Claude Fable 5 / Mythos 5 — 85.0% OSWorld-Verified, Anthropic-reported (Fable 5 and Mythos 5 both 85.0%). Kimi K3 (84.8%, Moonshot) and Claude Opus 4.8 (83.4%) sit just behind, all above Gemini 3.6 Flash's 83.0% — against a ~72% human baseline. All are vendor self-reports on the same task set.

In the timeline · 14 entries

More agents, tools & computer use benchmarks