Yutori

Last updated Aug 31, 2026

MacAgentBench Leaderboard

Tracking reported computer-use agent performance on MacAgentBench, combining the official benchmark leaderboard with additional results found by our Scout.

Reported results selected

ModelOrganizationAgentScore
Navigator n2Yutori83.1%1
Claude Opus 4.6AnthropicOpenClaw73.7%
Claude Opus 4.6AnthropicAgent-S366.9%
GPT-5.5OpenAI66.7%2
Gemini 3.1 ProGoogleOpenClaw63.3%
GPT-5.4OpenAIOpenClaw60.7%
GPT-5.4OpenAIAgent-S358.9%
GPT-5.4OpenAIBaseline58.4%
Claude Opus 4.8Anthropic58.4%2
Gemini 3.1 ProGoogleAgent-S354.3%
Qwen3VL-235B-A22BAlibabaAgent-S346.7%
Qwen3VL-235B-A22BAlibabaOpenClaw46.0%
Claude Opus 4.6AnthropicBaseline39.2%
Qwen3VL-32BAlibabaAgent-S337.4%
Gemini 3.1 ProGoogleBaseline34.2%
Qwen3VL-32BAlibabaOpenClaw33.7%
Qwen3VL-8BAlibabaAgent-S328.6%
Qwen3VL-8BAlibabaOpenClaw26.9%
Qwen3VL-235B-A22BAlibabaBaseline21.6%
Qwen3VL-32BAlibabaBaseline21.3%
OpenCUA-32BXLangBaseline18.8%
GUI-Owl-1.5-32BAlibabaBaseline17.2%
OpenCUA-7BXLangBaseline15.1%
Qwen3VL-8BAlibabaBaseline14.5%
InternVL3.5-14BOpenGVLabAgent-S313.5%
UI-TARS-72B-DPOByteDanceBaseline13.2%
ScaleCUA-32BOpenGVLabBaseline10.5%
GUI-Owl-1.5-8BAlibabaBaseline10.2%
UI-TARS-1.5-7BByteDanceBaseline9.8%
InternVL3.5-8BOpenGVLabAgent-S39.6%
ScaleCUA-7BOpenGVLabBaseline6.7%
InternVL3.5-14BOpenGVLabBaseline6.4%
InternVL3.5-8BOpenGVLabBaseline4.7%

Frequently asked questions

Resources

Citation
@misc{fu2026macagentbench,
  title={MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop},
  author={Yikun Fu and Bowen Fu and Zhenyu Wu and Shuang Cheng and Xiaowei Sun and Bowen Yang and Zehao Li and Yibo Zhao and Zichen Ding and Zhoumianze Liu and Shijie Wang and Biqing Qi and Bowen Zhou},
  year={2026},
  eprint={2606.22557},
  archivePrefix={arXiv},
  primaryClass={cs.AI},
}