# Online-Mind2Web Leaderboard

> Last updated AUG 05, 2026

A live ranking of web agents on 300 realistic tasks across 136 websites. Hark leads human evaluation at 97.7%; Yutori's Navigator n1.5 ranks #2 at 97.3%.

Data is reproduced from the official Online-Mind2Web leaderboard maintained by the OSU-NLP Group. Yutori Scouts monitor the source for changes.

## Human evaluation

Results listed on the official human-evaluation leaderboard are verified by the Online-Mind2Web team.

| Agent | Model | Organization | Source | Easy | Medium | Hard | Average SR | Date |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [Hark](<https://hark.com/>) | Hark Handoff | [Hark](<https://hark.com/>) | [Careerflow.ai](<https://www.careerflow.ai/human-data>) | 100.0 | 95.0 | 100.0 | 97.7% | AUG 04, 2026 |
| [Navigator](<https://yutori.com/blog/introducing-n1-5>) | n1.5-latest | [Yutori](<https://yutori.com>) | [Careerflow.ai](<https://www.careerflow.ai/human-data>) | 100.0 | 96.5 | 96.1 | 97.3% | JUN 28, 2026 |
| [ACT-2](<https://www.enhans.ai/act-2>) | GPT-5.4 | [Enhans](<https://www.enhans.ai/>) | [Careerflow.ai](<https://www.careerflow.ai/human-data>) | 93.0 | 94.3 | 89.9 | 92.7% | AUG 04, 2026 |
| [Navigator](<https://yutori.com/blog/introducing-navigator>) | n1-preview-11-2025 | [Yutori](<https://yutori.com>) | [Halluminate](<https://halluminate.ai/>) | 90.1 | 76.2 | 71.1 | 78.7% | NOV 18, 2025 |
| Google Computer Use (09-2025) | Gemini 2.5 Computer Use | Google DeepMind | Google DeepMind | 77.1 | 71.3 | 55.4 | 69.0% | SEP 29, 2025 |
| Operator | OpenAI Computer-Using Agent | OpenAI | OSU NLP | 83.1 | 58.0 | 43.2 | 61.3% | MAR 22, 2025 |
| ACT-1-20250814 | o3-2025-04-16 and Claude-sonnet-4-20250514 | Enhans | Enhans | 81.9 | 54.5 | 35.1 | 57.3% | AUG 23, 2025 |
| Claude Computer Use 3.7 (w/o thinking) | Claude-3-7-sonnet-20250219 | Anthropic | OSU NLP | 90.4 | 49.0 | 32.4 | 56.3% | APR 20, 2025 |
| ACT-1-20250703 | o3-2025-04-16 and Claude-sonnet-4-20250514 | Enhans | Enhans | 65.1 | 46.2 | 23.0 | 45.7% | JUL 16, 2025 |
| SeeAct | gpt-4o-2024-08-06 | OSU | OSU NLP | 60.2 | 25.2 | 8.1 | 30.7% | MAR 22, 2025 |
| Browser Use | gpt-4o-2024-08-06 | Browser Use | OSU NLP | 55.4 | 26.6 | 8.1 | 30.0% | MAR 22, 2025 |
| Claude Computer Use 3.5 | claude-3-5-sonnet-20241022 | Anthropic | OSU NLP | 56.6 | 20.3 | 14.9 | 29.0% | MAR 22, 2025 |
| Agent-E | gpt-4o-2024-08-06 | Emergence AI | OSU NLP | 49.4 | 26.6 | 6.8 | 28.0% | MAR 22, 2025 |

## Automatic evaluation

Automatic evaluation uses WebJudge with an o4-mini backbone. Unverified results are self-reported and their accuracy is not guaranteed.

| Agent | Model | Organization | Source | Easy | Medium | Hard | Average SR | Date | Verified | Note |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [Hark](<https://hark.com/>) | Hark Handoff | [Hark](<https://hark.com/>) | [Careerflow.ai](<https://www.careerflow.ai/human-data>) | 98.8 | 94.3 | 93.7 | 95.3% | AUG 04, 2026 | Yes | Include a concise final response in the judge; verified for factual consistency and without hallucinations. |
| [ACT-2](<https://www.enhans.ai/act-2>) | GPT-5.4 | [Enhans](<https://www.enhans.ai/>) | [Careerflow.ai](<https://www.careerflow.ai/human-data>) | 95.0 | 91.5 | 91.1 | 92.3% | AUG 04, 2026 | Yes | – |
| [Navigator](<https://yutori.com/blog/introducing-n1-5>) | n1.5-latest | [Yutori](<https://yutori.com>) | [Careerflow.ai](<https://www.careerflow.ai/>) | 97.5 | 84.4 | 84.4 | 87.9% | JUN 28, 2026 | Yes | Include a concise final response in the judge; verified for factual consistency and without hallucinations. |
| [Webwright](<https://www.microsoft.com/en-us/research/articles/webwright-a-terminal-is-all-you-need-for-web-agents/>) | GPT‑5.4 | [Microsoft](<https://www.microsoft.com/en-us/research/articles/webwright-a-terminal-is-all-you-need-for-web-agents/>) | [Microsoft](<https://www.microsoft.com/en-us/research/articles/webwright-a-terminal-is-all-you-need-for-web-agents/>) | 93.8 | 88.1 | 76.6 | 86.7% | MAY 04, 2026 | Yes | – |
| [Navigator](<https://yutori.com/blog/introducing-navigator>) | n1-preview-11-2025 | [Yutori](<https://yutori.com/>) | [Yutori](<https://yutori.com/>) | 84.0 | 62.2 | 48.7 | 64.7% | NOV 18, 2025 | Yes | – |
| Operator | OpenAI Computer-Using Agent | OpenAI | [OSU NLP](<https://arxiv.org/abs/2504.01382>) | 73.5 | 59.4 | 39.2 | 58.3% | MAY 11, 2025 | Yes | – |
| Google Computer Use (09-2025) | Gemini 2.5 Computer Use | Google DeepMind | [Google DeepMind](<https://blog.google/technology/google-deepmind/gemini-computer-use-model/>) | 77.1 | 55.2 | 45.9 | 57.3% | SEP 29, 2025 | Yes | – |
| ACT-1-20250814 | o3-2025-04-16 and Claude-sonnet-4-20250514 | Enhans | [Enhans](<https://www.enhans.ai/>) | 71.1 | 52.4 | 32.4 | 52.7% | AUG 23, 2025 | Yes | – |
| Claude Computer Use 3.7 (w/o thinking) | Claude-3-7-sonnet-20250219 | Anthropic | [OSU NLP](<https://arxiv.org/abs/2504.01382>) | 75.9 | 41.3 | 27.0 | 47.3% | MAY 11, 2025 | Yes | – |
| ACT-1-20250703 | o3-2025-04-16 and Claude-sonnet-4-20250514 | Enhans | [Enhans](<https://www.enhans.ai/>) | 53.7 | 39.2 | 24.3 | 39.5% | JUL 16, 2025 | Yes | – |
| [MolmoWeb](<https://allenai.org/blog/molmoweb>) | MolmoWeb-8B | [Allen AI](<https://allenai.org/>) | [Allen AI](<https://allenai.org/blog/molmoweb>) | 59.0 | 36.4 | 10.8 | 36.3% | MAR 24, 2026 | Yes | – |
| SeeAct | gpt-4o-2024-08-06 | OSU | [OSU NLP](<https://arxiv.org/abs/2504.01382>) | 51.8 | 28.0 | 9.5 | 30.0% | MAY 11, 2025 | Yes | – |
| Agent-E | gpt-4o-2024-08-06 | Emergence AI | [OSU NLP](<https://arxiv.org/abs/2504.01382>) | 51.8 | 23.1 | 6.8 | 27.0% | MAY 11, 2025 | Yes | – |
| Browser Use | gpt-4o-2024-08-06 | Browser Use | [OSU NLP](<https://arxiv.org/abs/2504.01382>) | 44.6 | 23.1 | 10.8 | 26.0% | MAY 11, 2025 | Yes | – |
| Claude Computer Use 3.5 | Claude-3-5-sonnet-20241022 | Anthropic | [OSU NLP](<https://arxiv.org/abs/2504.01382>) | 51.8 | 16.1 | 8.1 | 24.0% | MAY 11, 2025 | Yes | – |
| [GPT‑5.4](<https://openai.com/index/introducing-gpt-5-4/>) | GPT‑5.4 | [OpenAI](<https://openai.com/>) | [OpenAI](<https://openai.com/index/introducing-gpt-5-4/>) | – | – | – | 92.8% | MAR 05, 2026 | No | – |
| Eko-V2 | Unknown | Fellou | [Fellou](<https://fellou.ai/blog/post/eko20-launch/>) | 95.0 | 76.0 | 70.0 | 78.0% | MAY 24, 2025 | No | Unknown evaluation method |
| Seed1.5-VL | Seed1.5-VL | ByteDance | [ByteDance](<https://arxiv.org/pdf/2505.07062>) | – | – | – | 76.4% | MAY 11, 2025 | No | Evaluated by WebJudge(GPT-4o) |
| [ChatGPT Atlas](<https://openai.com/index/introducing-gpt-5-4/>) | ChatGPT Atlas | [OpenAI](<https://openai.com/>) | [OpenAI](<https://openai.com/index/introducing-gpt-5-4/>) | – | – | – | 70.9% | MAR 05, 2026 | No | – |
| Eko-V1 | Unknown | Fellou | [Fellou](<https://fellou.ai/blog/post/eko20-launch/>) | – | – | – | 31.0% | MAY 24, 2025 | No | Unknown evaluation method |

## About the benchmark

Online-Mind2Web evaluates web agents on 300 tasks across 136 live websites. Tasks are grouped by the number of steps a human annotator needs: Easy (1–5 steps), Medium (6–10 steps), and Hard (11+ steps). The primary metric is task success rate.

## Official sources

- [Official leaderboard](<https://huggingface.co/spaces/osunlp/Online_Mind2Web_Leaderboard>)
- [Paper](<https://arxiv.org/abs/2504.01382>)
- [Code](<https://github.com/OSU-NLP-Group/Online-Mind2Web>)
- [Dataset](<https://huggingface.co/datasets/osunlp/Online-Mind2Web>)

Source revision: [ba470716b95a67404059cb052710146ecec2bd68](<https://huggingface.co/spaces/osunlp/Online_Mind2Web_Leaderboard/commit/ba470716b95a67404059cb052710146ecec2bd68>)
