Currently scouting 6 benchmarks
Computer-use leaderboards
Yutori is a startup pushing the frontier of computer-use AI agents that can reliably act and execute long-running, multi-step tasks. To inform our mission, we rigorously track relevant benchmarks to keep track of state-of-the-art capabilities. To support the CUA community, we've made this data easily accessible here.
Leaderboards
Last updated Aug 07, 2026
OSWorld-Verified
The benchmark team’s corrected OSWorld release, with 369 tasks (361 without the Google Drive tasks requiring manual configuration), community-reported fixes, and maintainer-run results under consistent settings and environments.
Last updated Aug 05, 2026
Online-Mind2Web
300 real-world web tasks across 136 live sites and diverse domains. Tasks span Easy (1–5 steps), Medium (6–10), and Hard (11+), with WebJudge automatic evaluation.
Last updated Jul 22, 2026
WeaveBench
114 long-horizon tasks across 8 work domains, each requiring GUI interaction plus CLI or code execution within the same trajectory.
Last updated Jul 15, 2026
OSWorld 2.0
108 long-horizon workflows in real computer environments, extending desktop-agent evaluation beyond short, isolated actions into stateful tasks with many executable checkpoints.
Last updated Jul 15, 2026
MyPCBench
Personal-assistant tasks on a desktop with coherent user identity, history, and logged-in accounts: 184 tasks across 17 web applications and the surrounding desktop stack.
Last updated Jun 21, 2026
MacAgentBench
676 tasks across 25 real macOS applications, including single- and multi-application workflows with GUI and CLI interaction.