Last updated Aug 31, 2026
MacAgentBench Leaderboard
Tracking reported computer-use agent performance on MacAgentBench, combining the official benchmark leaderboard with additional results found by our Scout.
Reported results selected
| Model | Organization | Agent | Score |
|---|---|---|---|
| Navigator n2 | Yutori | – | 83.1%1 |
| Claude Opus 4.6 | Anthropic | OpenClaw | 73.7% |
| Claude Opus 4.6 | Anthropic | Agent-S3 | 66.9% |
| GPT-5.5 | OpenAI | – | 66.7%2 |
| Gemini 3.1 Pro | OpenClaw | 63.3% | |
| GPT-5.4 | OpenAI | OpenClaw | 60.7% |
| GPT-5.4 | OpenAI | Agent-S3 | 58.9% |
| GPT-5.4 | OpenAI | Baseline | 58.4% |
| Claude Opus 4.8 | Anthropic | – | 58.4%2 |
| Gemini 3.1 Pro | Agent-S3 | 54.3% | |
| Qwen3VL-235B-A22B | Alibaba | Agent-S3 | 46.7% |
| Qwen3VL-235B-A22B | Alibaba | OpenClaw | 46.0% |
| Claude Opus 4.6 | Anthropic | Baseline | 39.2% |
| Qwen3VL-32B | Alibaba | Agent-S3 | 37.4% |
| Gemini 3.1 Pro | Baseline | 34.2% | |
| Qwen3VL-32B | Alibaba | OpenClaw | 33.7% |
| Qwen3VL-8B | Alibaba | Agent-S3 | 28.6% |
| Qwen3VL-8B | Alibaba | OpenClaw | 26.9% |
| Qwen3VL-235B-A22B | Alibaba | Baseline | 21.6% |
| Qwen3VL-32B | Alibaba | Baseline | 21.3% |
| OpenCUA-32B | XLang | Baseline | 18.8% |
| GUI-Owl-1.5-32B | Alibaba | Baseline | 17.2% |
| OpenCUA-7B | XLang | Baseline | 15.1% |
| Qwen3VL-8B | Alibaba | Baseline | 14.5% |
| InternVL3.5-14B | OpenGVLab | Agent-S3 | 13.5% |
| UI-TARS-72B-DPO | ByteDance | Baseline | 13.2% |
| ScaleCUA-32B | OpenGVLab | Baseline | 10.5% |
| GUI-Owl-1.5-8B | Alibaba | Baseline | 10.2% |
| UI-TARS-1.5-7B | ByteDance | Baseline | 9.8% |
| InternVL3.5-8B | OpenGVLab | Agent-S3 | 9.6% |
| ScaleCUA-7B | OpenGVLab | Baseline | 6.7% |
| InternVL3.5-14B | OpenGVLab | Baseline | 6.4% |
| InternVL3.5-8B | OpenGVLab | Baseline | 4.7% |
1 Navigator n2’s MacAgentBench score is reported in the launch post.
2 Claude Opus 4.8 and GPT-5.5 results reported by Alibaba in the Qwen-CUA technical report.
Frequently asked questions
Resources
Paper (arXiv 2606.22557)Code (GitHub · JetAstra)Environment (Hugging Face)Official leaderboardNavigator n2 launch post
Citation
@misc{fu2026macagentbench,
title={MacAgentBench: Benchmarking AI Agents on Real-World macOS Desktop},
author={Yikun Fu and Bowen Fu and Zhenyu Wu and Shuang Cheng and Xiaowei Sun and Bowen Yang and Zehao Li and Yibo Zhao and Zichen Ding and Zhoumianze Liu and Shijie Wang and Biqing Qi and Bowen Zhou},
year={2026},
eprint={2606.22557},
archivePrefix={arXiv},
primaryClass={cs.AI},
}