Last updated Jul 15, 2026
MyPCBench Leaderboard
The official leaderboard's latest evaluation data, as reported by our Scout
Live leaderboard selected
| Model | Vendor | Weights | Perfect | Rubric score | Perfect tasks | Max steps | Avg. steps | Traj. efficiency | Agent interface | Self-reported |
|---|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 4.8 | Anthropic | Closed | 62.0% | 88.8% | 114 | 200 | – | – | Computer only | Yes |
| Qwen-CUA | Qwen × XLang | Unreleased | 58.7% | 84.3% | 108 | 200 | – | – | Computer only | Yes |
| Claude Opus 4.6 | Anthropic | Closed | 58.2% | 85.1% | 107 | 100 | 46.5 | 3.78 | Computer + Bash | No |
| GPT-5.6-sol | OpenAI | Closed | 55.4% | 85.3% | 102 | 100 | 57.2 | 2.26 | Computer + Bash | No |
| GPT-5.6-luna | OpenAI | Closed | 55.4% | 83.2% | 102 | 100 | 63.5 | 2.00 | Computer + Bash | No |
| Qwen-3.7 | Qwen | Unreleased | 51.6% | 81.5% | 95 | 200 | – | – | Computer only | Yes |
| Claude Sonnet 4.6 | Anthropic | Closed | 50.5% | 78.5% | 93 | 100 | 45.8 | 3.68 | Computer + Bash | No |
| GPT-5.5 | OpenAI | Closed | 45.1% | 83.6% | 83 | 100 | 43.5 | 2.80 | Computer + Bash | No |
| GPT-5.4 mini | OpenAI | Closed | 23.9% | 53.5% | 44 | 100 | 49.5 | 1.64 | Computer + Bash | No |
| Qwen 3.5 35B-A3B | Alibaba | Open | 7.6% | 42.4% | 14 | 100 | 66.0 | 1.41 | Computer + Bash | No |
| Qwen 3.5 9B | Alibaba | Open | 6.5% | 24.1% | 12 | 100 | 69.2 | 1.09 | Computer + Bash | No |
Frequently asked questions
Resources
Citation
@misc{jang2026mypcbench,
title={MyPCBench: A Benchmark for Personally Intelligent Computer-Use Agents},
author={Jang, Lawrence Keunho and Jang, Andrew Keunwoo and Koh, Jing Yu and Salakhutdinov, Ruslan},
year={2026},
eprint={2606.16748},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://mypcbench.com},
}