OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
- Published
- Source
- arXiv
- Paper number
- 531
- Field
- AI / General
- arXiv ID
- 2606.29537
Key points
- It contains 108 long-horizon workflows, with a human median of 1.6 hours and an average of 318 tool calls per agent.
- Claude Opus 4.8 with max thinking completes 20.6 percent of tasks, while GPT-5.5 completes about 13 percent, although its token efficiency is better.
- It defines 10 challenge phenomena, including dynamic environments, cross-source reasoning, hidden-state reasoning, and conflict resolution.
- It builds a realistic environment from 31 self-hosted web services, including email, banking, and chat.
- The main failure causes are forgetting constraints, missing intermediate information, guessing, and skipping verification, not basic GUI manipulation.
- Agents spend less than 7 percent of their own budget on error detection and correction.
Paper links
External research summaries. These are not HDATF publications or measured product results.