# Computer-use leaderboards

Current official results across web and computer-use agent benchmarks. Open a leaderboard for its complete ranking, methodology notes, resources, citation, and machine-readable Markdown view.

| Benchmark | Description | Scope | Track | Current leader | Score | Updated |
| --- | --- | --- | --- | --- | --- | --- |
| [OSWorld-Verified](<https://yutori.com/leaderboards/osworld-verified>) | The benchmark team’s corrected OSWorld release, with 369 tasks (361 without the Google Drive tasks requiring manual configuration), community-reported fixes, and maintainer-run results under consistent settings and environments. | 369 tasks, 361 without Drive | Foundation E2E GUI | claude-fable-5[1m] (Anthropic) | 86.0% success rate | Aug 07, 2026 |
| [Online-Mind2Web](<https://yutori.com/leaderboards/online-mind2web>) | 300 real-world web tasks across 136 live sites and diverse domains. Tasks span Easy (1–5 steps), Medium (6–10), and Hard (11+), with WebJudge automatic evaluation. | 300 tasks across 136 websites | Human evaluation | Hark (Hark) | 97.7% success rate | Aug 05, 2026 |
| [WeaveBench](<https://yutori.com/leaderboards/weavebench>) | 114 long-horizon tasks across 8 work domains, each requiring GUI interaction plus CLI or code execution within the same trajectory. | 114 tasks across 8 domains | Live leaderboard | Claude Opus 4.7 (Claude Code) | 41.2% PassRate | Jul 22, 2026 |
| [OSWorld 2.0](<https://yutori.com/leaderboards/osworld-2>) | 108 long-horizon workflows in real computer environments, extending desktop-agent evaluation beyond short, isolated actions into stateful tasks with many executable checkpoints. | 108 long-horizon workflows | 500-step track | Claude Opus 4.8 (Anthropic) | 20.60% binary accuracy | Jul 15, 2026 |
| [MyPCBench](<https://yutori.com/leaderboards/mypcbench>) | Personal-assistant tasks on a desktop with coherent user identity, history, and logged-in accounts: 184 tasks across 17 web applications and the surrounding desktop stack. | 184 tasks across 17 applications | Self-reported | Claude Opus 4.8 (Anthropic) | 62.0% Perfect | Jul 15, 2026 |
| [MacAgentBench](<https://yutori.com/leaderboards/mac-agent-bench>) | 676 tasks across 25 real macOS applications, including single- and multi-application workflows with GUI and CLI interaction. | 676 tasks across 25 applications | Official Pass@1 | Claude Opus 4.6 (OpenClaw) | 73.7% Pass@1 | Jun 21, 2026 |

Scores from different environments, step budgets, and evaluation methods are not directly comparable.
