OpenCUA: Open Foundations for Computer-Use Agents

Published
Source
arXiv
Paper number
076
Field
Computer Use Agents
arXiv ID
2508.09123

Key points

  • Most state-of-the-art computer-use agent, or CUA, systems are proprietary and closed-source, which makes it harder to study their capabilities, limitations, and risks.
  • There is no scalable open-source infrastructure for collecting diverse, large-scale computer-use data, including real-time user interaction and accessibility trees.
  • Existing open-source CUA datasets have limited scope, scale, and diversity, and they focus on narrow domains rather than general computer-use applications.
  • The authors developed AGENTNET TOOL, a cross-platform application that captures human computer-use demonstrations, including screen video, mouse and keyboard signals, and accessibility trees, across Windows, macOS, and Ubuntu.
  • They built AGENTNET, the first large-scale and diverse computer-use task dataset with 22,625 trajectories spanning multiple operating systems and more than 200 applications and websites.
  • To improve agent planning and error recovery, they propose a scalable training pipeline that integrates a new reflective long-chain-of-thought synthesis and multi-image history context encoding.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)