# OSWorld 2.0 Leaderboard

> Last updated JUL 15, 2026

A current ranking of computer-use agents on 108 long-horizon tasks. Claude Opus 4.8 leads the official 500-step track at 20.6% binary accuracy.

Data is reproduced from the official OSWorld 2.0 leaderboard maintained by the XLANG Lab. Yutori Scouts monitor the source for changes.

## 150 steps

| Model | Organization | Reasoning | Tool setting | Binary accuracy | Partial score | Est. cost |
| --- | --- | --- | --- | --- | --- | --- |
| GPT-5.5 | OpenAI | XHigh | Batch Tool | 13.00% | 46.70% | $1,875 |
| Claude Opus 4.7 | Anthropic | Max | Standard | 4.60% | 20.30% | $1,030 |
| Claude Sonnet 4.6 | Anthropic | Max | Standard | 4.60% | 20.00% | $800 |
| Claude Sonnet 4.6 | Anthropic | Medium | Standard | 4.60% | 14.20% | $410 |
| MiniMax M3 | MiniMax | Enabled | Standard | 1.90% | 8.20% | $87 |
| Kimi 2.6 | Moonshot AI | Enabled | Standard | 1.90% | 7.10% | $336 |

## 300 steps

| Model | Organization | Reasoning | Tool setting | Binary accuracy | Partial score | Est. cost |
| --- | --- | --- | --- | --- | --- | --- |
| GPT-5.5 | OpenAI | XHigh | Batch Tool | 13.00% | 49.50% | $2,750 |
| Claude Opus 4.7 | Anthropic | Max | Standard | 13.00% | 39.80% | $2,470 |
| Claude Sonnet 4.6 | Anthropic | Medium | Standard | 8.30% | 29.40% | $990 |
| Claude Sonnet 4.6 | Anthropic | Max | Standard | 6.50% | 35.80% | $1,720 |
| Kimi 2.6 | Moonshot AI | Enabled | Standard | 4.60% | 14.40% | $604 |
| MiniMax M3 | MiniMax | Enabled | Standard | 3.70% | 16.60% | $182 |
| Qwen 3.7-Plus | Alibaba | Thinking | Standard | 1.90% | 16.60% | $403 |

## 500 steps

| Model | Organization | Reasoning | Tool setting | Binary accuracy | Partial score | Est. cost |
| --- | --- | --- | --- | --- | --- | --- |
| Claude Opus 4.8 | Anthropic | Max | Batched Tool | 20.60% | 54.80% | – |
| Claude Opus 4.8 | Anthropic | Max | Standard | 18.52% | 49.33% | – |
| Claude Opus 4.7 | Anthropic | Max | Batched Tool | 18.20% | 48.91% | – |
| Claude Opus 4.7 | Anthropic | Max | Standard | 13.90% | 49.10% | $3,870 |
| GPT-5.5 | OpenAI | XHigh | Batch Tool | 13.00% | 49.50% | $2,750 |
| Claude Sonnet 4.6 | Anthropic | Medium | Standard | 9.30% | 33.90% | $1,550 |
| Claude Sonnet 4.6 | Anthropic | Max | Standard | 8.30% | 41.50% | $2,410 |
| MiniMax M3 | MiniMax | Enabled | Standard | 4.60% | 22.30% | $259 |
| Kimi 2.6 | Moonshot AI | Enabled | Standard | 4.60% | 22.10% | $708 |
| Qwen 3.7-Plus | Alibaba | Thinking | Standard | 2.80% | 21.50% | $412 |

## About the benchmark

OSWorld 2.0 evaluates long-horizon computer-use workflows across self-hosted web environments and professional desktop applications. Each task contains multiple executable checkpoints, supporting both binary completion accuracy and partial-credit scoring.

## Official sources

- [Paper (arXiv 2606.29537)](<https://arxiv.org/abs/2606.29537>)
- [Code (GitHub · XLANG)](<https://github.com/xlang-ai/OSWorld-V2>)
- [Dataset (Hugging Face)](<https://huggingface.co/datasets/xlangai/osworld_v2_tasks>)
- [Official leaderboard](<https://osworld-v2.xlang.ai/#leaderboard>)

Source revision (SHA-256): `a48d24c886c62a64336dc827c5b3cb46c61822de38522785965f804ba76f4aaf`
