# MyPCBench Leaderboard

> Last updated JUL 15, 2026

A current ranking of personal computer-use agents on 184 tasks across 17 applications. Claude Opus 4.8 leads at 62.0% Perfect.

Data is reproduced from the official MyPCBench leaderboard maintained by the benchmark authors at Carnegie Mellon University. Yutori Scouts monitor the source for changes.

## Live leaderboard

| Model | Vendor | Weights | Perfect | Rubric score | Perfect tasks | Max steps | Avg. steps | Traj. efficiency | Agent interface | Self-reported |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [Claude Opus 4.8](<https://github.com/xlang-ai/Qwen-CUA/blob/main/paper/Qwen-CUA.pdf>) | Anthropic | Closed | 62.0% | 88.8% | 114 | 200 | – | – | Computer only | Yes |
| [Qwen-CUA](<https://github.com/xlang-ai/Qwen-CUA/blob/main/paper/Qwen-CUA.pdf>) | Qwen × XLang | Unreleased | 58.7% | 84.3% | 108 | 200 | – | – | Computer only | Yes |
| Claude Opus 4.6 | Anthropic | Closed | 58.2% | 85.1% | 107 | 100 | 46.5 | 3.78 | Computer + Bash | No |
| GPT-5.6-sol | OpenAI | Closed | 55.4% | 85.3% | 102 | 100 | 57.2 | 2.26 | Computer + Bash | No |
| GPT-5.6-luna | OpenAI | Closed | 55.4% | 83.2% | 102 | 100 | 63.5 | 2.00 | Computer + Bash | No |
| [Qwen-3.7](<https://github.com/xlang-ai/Qwen-CUA/blob/main/paper/Qwen-CUA.pdf>) | Qwen | Unreleased | 51.6% | 81.5% | 95 | 200 | – | – | Computer only | Yes |
| Claude Sonnet 4.6 | Anthropic | Closed | 50.5% | 78.5% | 93 | 100 | 45.8 | 3.68 | Computer + Bash | No |
| GPT-5.5 | OpenAI | Closed | 45.1% | 83.6% | 83 | 100 | 43.5 | 2.80 | Computer + Bash | No |
| GPT-5.4 mini | OpenAI | Closed | 23.9% | 53.5% | 44 | 100 | 49.5 | 1.64 | Computer + Bash | No |
| Qwen 3.5 35B-A3B | Alibaba | Open | 7.6% | 42.4% | 14 | 100 | 66.0 | 1.41 | Computer + Bash | No |
| Qwen 3.5 9B | Alibaba | Open | 6.5% | 24.1% | 12 | 100 | 69.2 | 1.09 | Computer + Bash | No |

## About the benchmark

MyPCBench evaluates agents on a reproducible Linux desktop seeded with one user’s identity, history, and logged-in accounts. Its 184 tasks span 17 web applications and the desktop stack, with 1,191 natural-language rubric items.

## Official sources

- [Paper (arXiv 2606.16748)](<https://arxiv.org/abs/2606.16748>)
- [Code (GitHub)](<https://github.com/ljang0/MyPCBench>)
- [Dataset (Hugging Face)](<https://huggingface.co/datasets/ljang0/mypcbench-qemu-baseline>)
- [Official leaderboard](<https://mypcbench.com/leaderboard>)

Source revision (SHA-256): `fc11fde8b4564f3567f1127fd7be6183dfacffb2b1f7fd0a5192e173eca94a93`
