| Rank | Model | Score |
|---|---|---|
| 1 | kimi-k3 | 58.3 |
1 phaseActive
Expanded, harder revision of OSWorld's real computer-use agent benchmark — broader app coverage and tougher multi-step tasks. Metric: % success.
Quick answer: OSWorld 2.0 is an expanded, more difficult revision of the original OSWorld computer-use agent benchmark, broadening application coverage and task complexity beyond the original 369-task suite. Kimi K3 scores 58.3%, notably lower than its 84.8% on OSWorld-Verified — reflecting the harder task pool introduced in the 2.0 revision.
What it tests: Real desktop and web GUI task completion, with a broader and harder task pool than the original OSWorld release — including longer workflows and more diverse applications.
Why it matters: As agents saturate the original OSWorld task suite, OSWorld 2.0 provides continued headroom to differentiate frontier computer-use agents.
Known limitations: As a newer revision, less third-party leaderboard data exists compared to the original OSWorld/OSWorld-Verified suite, so cross-model comparisons are still limited.
OSWorld 2.0 builds on the original OSWorld environment — an interactive, scalable real-computer sandbox spanning Ubuntu, Windows, and macOS — with an expanded and more challenging task pool. Tasks continue to use execution-based evaluation (checking final system/application state) rather than LLM-as-judge scoring, preserving OSWorld's reproducibility advantage while raising the difficulty ceiling as frontier agents improve.
| Field | Value |
|---|---|
| Task category | Agent / computer use (GUI) |
| Metric | % task success (execution-based) |
| Operating systems | Ubuntu, Windows, macOS |
| Saturation | Low |
| Based on | OSWorld (Xie et al., arXiv 2404.07972) |
| GitHub | xlang-ai/OSWorld |
Tasks are scored pass/fail via execution-based scripts that verify the resulting OS or application state matches the expected outcome. Aggregate scores are reported as % of tasks completed successfully.
| Rank | Model | Score | Source | Date |
|---|---|---|---|---|
| 1 | Kimi K3 | 58.3% | Kimi K3 technical report | 2026-07 |
Score sourced from Moonshot AI's Kimi K3 announcement, July 2026.
No Benchgen results yet — be the first to run OSWorld 2.0.
| Benchmark | What it tests | Saturation |
|---|---|---|
| OSWorld 2.0 | Expanded, harder GUI agent workflows | Low |
| OSWorld-Verified | Original, corrected GUI agent task suite | Low |
| MCP Atlas | MCP tool-use agentic workflows | Low |
| SaaS-Bench | SaaS product usage workflows | Low |
Benchgen lets you run OSWorld 2.0 against your own model and harness configuration, tracking GUI task success over time as you iterate on agent scaffolding and fine-tuning.