Last updated Aug 31, 2026
WeaveBench Leaderboard
Tracking reported computer-use agent performance on WeaveBench, combining the official benchmark leaderboard with additional results found by our Scout.
Reported results selected
| Model | Harness | Overall score | PR |
|---|---|---|---|
| Yutori Navigator n21 | Yutori Agent | 70.3% | 54.75%2 |
| Seed 2.1 turbo | Claude Code | 57.1% | 28.9% |
| Seed 2.1 pro | Claude Code | 55.1% | 30.7% |
| Claude Opus 4.7 | Claude Code | 53.2% | 41.2% |
| Claude Opus 4.7 | Hermes Agent | 51.6% | 28.1% |
| GPT-5.5 | Codex CLI | 49.9% | 35.1% |
| Claude Opus 4.7 | OpenClaw | 48.2% | 35.1% |
| GPT-5.5 | OpenClaw | 46.6% | 33.3% |
| GPT-5.5 | Hermes Agent | 46.6% | 31.6% |
| GPT-5.4 | OpenClaw | 46.5% | 22.8% |
| GPT-5.3-codex | OpenClaw | 45.6% | 18.4% |
| Claude Opus 4.7 | Codex CLI | 37.8% | 13.2% |
| GPT-5.2-codex | OpenClaw | 32.1% | 6.1% |
| Qwen3.5-397B-A17B | OpenClaw | 31.8% | 0.9% |
| GPT-5.5 | Claude Code | 29.9% | 14.9% |
| GPT-5.1-codex | OpenClaw | 22.6% | 1.8% |
| Gemini 3.1 Pro | OpenClaw | 22.3% | 1.8% |
| Qwen3-VL-8B-Think | OpenClaw | 9.2% | 0.9% |
| GUI-Owl-1.5-32B | OpenClaw | 6.5% | 0.0% |
1 Navigator n2’s Overall score is reported in the launch post.
2 Navigator n2’s 54.75% PR and Yutori Agent harness are first-party Yutori data supplied for this leaderboard. The PR was obtained with a 500-step limit.
Frequently asked questions
Resources
Paper (arXiv 2606.09426)Code (GitHub)Dataset (Hugging Face)Official leaderboardNavigator n2 launch post
Citation
@article{li2026weavebench,
title={WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces},
author={Li, Wanli and Zhou, Bowen and Yu, Yunyao and Xu, Zhou and Yang, Yifan and Li, Dongsheng and Shan, Caihua},
year={2026},
eprint={2606.09426},
archivePrefix={arXiv},
primaryClass={cs.AI},
}