Last updated Jul 22, 2026
WeaveBench Leaderboard
The official leaderboard's latest evaluation data, as reported by our Scout
Live leaderboard selected
| Model | Harness | PR | Overall | DSK | DOC | GAM | WEB | DAV | OPS | SPA | DES |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 4.7 | Claude Code | 41.2% | 0.532 | 55.6% | 47.1% | 23.5% | 53.3% | 23.1% | 50.0% | 33.3% | 40.0% |
| GPT-5.5 | Codex CLI | 35.1% | 0.499 | 38.9% | 29.4% | 23.5% | 53.3% | 15.4% | 50.0% | 58.3% | 10.0% |
| Claude Opus 4.7 | OpenClaw | 35.1% | 0.482 | 55.6% | 29.4% | 23.5% | 66.7% | 15.4% | 41.7% | 16.7% | 20.0% |
| GPT-5.5 | OpenClaw | 33.3% | 0.466 | 38.9% | 35.3% | 35.3% | 21.4% | 23.1% | 38.5% | 33.3% | 40.0% |
| GPT-5.5 | Hermes Agent | 31.6% | 0.466 | 55.6% | 29.4% | 35.3% | 40.0% | 7.7% | 25.0% | 25.0% | 20.0% |
| Seed 2.1 pro | Claude Code | 30.7% | 0.551 | 44.4% | 41.2% | 17.6% | 40.0% | 30.8% | 8.3% | 33.3% | 20.0% |
| Seed 2.1 turbo | Claude Code | 28.9% | 0.571 | 38.9% | 35.3% | 17.6% | 33.3% | 30.8% | 16.7% | 33.3% | 20.0% |
| Claude Opus 4.7 | Hermes Agent | 28.1% | 0.516 | 33.3% | 47.1% | 11.8% | 26.7% | 30.8% | 50.0% | 8.3% | 10.0% |
| GPT-5.4 | OpenClaw | 22.8% | 0.465 | 55.6% | 35.3% | 5.9% | 0.0% | 23.1% | 23.1% | 8.3% | 20.0% |
| GPT-5.3-codex | OpenClaw | 18.4% | 0.456 | 33.3% | 23.5% | 29.4% | 0.0% | 7.7% | 16.7% | 8.3% | 20.0% |
| GPT-5.5 | Claude Code | 14.9% | 0.299 | 33.3% | 11.8% | 11.8% | 0.0% | 15.4% | 16.7% | 25.0% | 0.0% |
| Claude Opus 4.7 | Codex CLI | 13.2% | 0.378 | 16.7% | 11.8% | 11.8% | 6.7% | 7.7% | 25.0% | 16.7% | 10.0% |
| GPT-5.2-codex | OpenClaw | 6.1% | 0.321 | 5.6% | 11.8% | 0.0% | 0.0% | 15.4% | 16.7% | 0.0% | 0.0% |
| GPT-5.1-codex | OpenClaw | 1.8% | 0.226 | 0.0% | 5.9% | 0.0% | 0.0% | 7.7% | 0.0% | 0.0% | 0.0% |
| Gemini 3.1 Pro | OpenClaw | 1.8% | 0.223 | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 8.3% | 8.3% | 0.0% |
| Qwen3.5-397B-A17B | OpenClaw | 0.9% | 0.318 | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 8.3% | 0.0% | 0.0% |
| Qwen3-VL-8B-Think | OpenClaw | 0.9% | 0.092 | 0.0% | 0.0% | 0.0% | 0.0% | 8.3% | 0.0% | 0.0% | 0.0% |
| GUI-Owl-1.5-32B | OpenClaw | 0.0% | 0.065 | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% | 0.0% |
Frequently asked questions
Resources
Citation
@article{li2026weavebench,
title={WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces},
author={Li, Wanli and Zhou, Bowen and Yu, Yunyao and Xu, Zhou and Yang, Yifan and Li, Dongsheng and Shan, Caihua},
year={2026},
eprint={2606.09426},
archivePrefix={arXiv},
primaryClass={cs.AI},
}