# MacAgentBench Leaderboard

> Last updated JUN 21, 2026

A current ranking of computer-use agents on 676 macOS tasks across 25 applications. Claude Opus 4.6 with OpenClaw leads at 73.7% Pass@1.

Data is reproduced from the official MacAgentBench project leaderboard. Yutori Scouts monitor the source for changes.

## Live leaderboard

| Model | Organization | Agent | Score | Steps | Tokens (in / out) | Cost / task |
| --- | --- | --- | --- | --- | --- | --- |
| Claude Opus 4.6 | Anthropic | OpenClaw | 73.7% | 8.6 | 263K / 1.2K | $1.35 |
| Claude Opus 4.6 | Anthropic | Agent-S3 | 66.9% | 10.0 | 262K / 4.4K | $1.30 |
| Gemini 3.1 Pro | Google | OpenClaw | 63.3% | 6.6 | 122K / 2.0K | $0.27 |
| GPT-5.4 | OpenAI | OpenClaw | 60.7% | 6.5 | 153K / 1.5K | $0.40 |
| GPT-5.4 | OpenAI | Agent-S3 | 58.9% | 10.2 | 263K / 3.0K | $0.66 |
| GPT-5.4 | OpenAI | Baseline | 58.4% | 13.2 | 127K / 3.0K | $0.36 |
| Gemini 3.1 Pro | Google | Agent-S3 | 54.3% | 8.1 | 136K / 8.7K | $0.36 |
| Qwen3VL-235B-A22B | Alibaba | Agent-S3 | 46.7% | 13.3 | 379K / 16.8K | $0.14 |
| Qwen3VL-235B-A22B | Alibaba | OpenClaw | 46.0% | 8.4 | 211K / 5.7K | $0.07 |
| Claude Opus 4.6 | Anthropic | Baseline | 39.2% | 27.1 | 177K / 4.3K | $0.98 |
| Qwen3VL-32B | Alibaba | Agent-S3 | 37.4% | 13.4 | 379K / 20.5K | $0.05 |
| Gemini 3.1 Pro | Google | Baseline | 34.2% | 26.5 | 127K / 19.7K | $0.49 |
| Qwen3VL-32B | Alibaba | OpenClaw | 33.7% | 8.3 | 209K / 7.6K | $0.02 |
| Qwen3VL-8B | Alibaba | Agent-S3 | 28.6% | 13.7 | 426K / 29.3K | $0.05 |
| Qwen3VL-8B | Alibaba | OpenClaw | 26.9% | 13.2 | 455K / 12.7K | $0.07 |
| Qwen3VL-235B-A22B | Alibaba | Baseline | 21.6% | 23.2 | 257K / 4.4K | $0.08 |
| Qwen3VL-32B | Alibaba | Baseline | 21.3% | 32.7 | 385K / 10.4K | $0.04 |
| OpenCUA-32B | XLang | Baseline | 18.8% | 20.1 | 176K / 8.3K | -- |
| GUI-Owl-1.5-32B | Alibaba | Baseline | 17.2% | 27.6 | 323K / 5.3K | -- |
| OpenCUA-7B | XLang | Baseline | 15.1% | 24.1 | 214K / 10.1K | -- |
| Qwen3VL-8B | Alibaba | Baseline | 14.5% | 35.2 | 468K / 35.9K | $0.10 |
| InternVL3.5-14B | OpenGVLab | Agent-S3 | 13.5% | 13.6 | 515K / 5.1K | -- |
| UI-TARS-72B-DPO | ByteDance | Baseline | 13.2% | 29.2 | 432K / 2.6K | -- |
| ScaleCUA-32B | OpenGVLab | Baseline | 10.5% | 33.5 | 135K / 4.1K | -- |
| GUI-Owl-1.5-8B | Alibaba | Baseline | 10.2% | 24.8 | 284K / 9.3K | -- |
| UI-TARS-1.5-7B | ByteDance | Baseline | 9.8% | 42.2 | 651K / 4.1K | $0.07 |
| InternVL3.5-8B | OpenGVLab | Agent-S3 | 9.6% | 21.6 | 1075K / 12.3K | -- |
| ScaleCUA-7B | OpenGVLab | Baseline | 6.7% | 30.3 | 121K / 4.3K | -- |
| InternVL3.5-14B | OpenGVLab | Baseline | 6.4% | 32.9 | 119K / 5.2K | -- |
| InternVL3.5-8B | OpenGVLab | Baseline | 4.7% | 32.1 | 115K / 4.7K | -- |

## About the benchmark

MacAgentBench is a macOS desktop-agent benchmark with 676 tasks across 25 applications. It uses deterministic rule-based evaluation, fine-grained checkpoints, and model comparisons across Baseline, Agent-S3, and OpenClaw frameworks.

## Official sources

- [Paper (arXiv 2606.22557)](<https://arxiv.org/abs/2606.22557>)
- [Code (GitHub · JetAstra)](<https://github.com/JetAstra/MacAgentBench>)
- [Environment (Hugging Face)](<https://huggingface.co/JetLM/OpenClaw-macOS>)
- [Official leaderboard](<https://jetastra.github.io/MacAgentBench/>)

Source revision (SHA-256): `20d9e2f8ede68266a3e60f2f20c4ff5644a4bb06bd9a46af8688413c7b3ca9e0`
