Yutori

Last updated Aug 31, 2026

WeaveBench Leaderboard

Tracking reported computer-use agent performance on WeaveBench, combining the official benchmark leaderboard with additional results found by our Scout.

Reported results selected

ModelHarnessOverall scorePR
Yutori Navigator n21Yutori Agent70.3%54.75%2
Seed 2.1 turboClaude Code57.1%28.9%
Seed 2.1 proClaude Code55.1%30.7%
Claude Opus 4.7Claude Code53.2%41.2%
Claude Opus 4.7Hermes Agent51.6%28.1%
GPT-5.5Codex CLI49.9%35.1%
Claude Opus 4.7OpenClaw48.2%35.1%
GPT-5.5OpenClaw46.6%33.3%
GPT-5.5Hermes Agent46.6%31.6%
GPT-5.4OpenClaw46.5%22.8%
GPT-5.3-codexOpenClaw45.6%18.4%
Claude Opus 4.7Codex CLI37.8%13.2%
GPT-5.2-codexOpenClaw32.1%6.1%
Qwen3.5-397B-A17BOpenClaw31.8%0.9%
GPT-5.5Claude Code29.9%14.9%
GPT-5.1-codexOpenClaw22.6%1.8%
Gemini 3.1 ProOpenClaw22.3%1.8%
Qwen3-VL-8B-ThinkOpenClaw9.2%0.9%
GUI-Owl-1.5-32BOpenClaw6.5%0.0%

1 Navigator n2’s Overall score is reported in the launch post.

2 Navigator n2’s 54.75% PR and Yutori Agent harness are first-party Yutori data supplied for this leaderboard. The PR was obtained with a 500-step limit.

Frequently asked questions

Resources

Citation
@article{li2026weavebench,
  title={WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces},
  author={Li, Wanli and Zhou, Bowen and Yu, Yunyao and Xu, Zhou and Yang, Yifan and Li, Dongsheng and Shan, Caihua},
  year={2026},
  eprint={2606.09426},
  archivePrefix={arXiv},
  primaryClass={cs.AI},
}