OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

Published
Source
arXiv
Paper number
531
Field
AI / General
arXiv ID
2606.29537

Key points

  • It contains 108 long-horizon workflows, with a human median of 1.6 hours and an average of 318 tool calls per agent.
  • Claude Opus 4.8 with max thinking completes 20.6 percent of tasks, while GPT-5.5 completes about 13 percent, although its token efficiency is better.
  • It defines 10 challenge phenomena, including dynamic environments, cross-source reasoning, hidden-state reasoning, and conflict resolution.
  • It builds a realistic environment from 31 self-hosted web services, including email, banking, and chat.
  • The main failure causes are forgetting constraints, missing intermediate information, guessing, and skipping verification, not basic GUI manipulation.
  • Agents spend less than 7 percent of their own budget on error detection and correction.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)