OpenCUA: Open Foundations for Computer-Use Agents
- Published
- Source
- arXiv
- Paper number
- 076
- Field
- Computer Use Agents
- arXiv ID
- 2508.09123
Key points
- Most state-of-the-art computer-use agent, or CUA, systems are proprietary and closed-source, which makes it harder to study their capabilities, limitations, and risks.
- There is no scalable open-source infrastructure for collecting diverse, large-scale computer-use data, including real-time user interaction and accessibility trees.
- Existing open-source CUA datasets have limited scope, scale, and diversity, and they focus on narrow domains rather than general computer-use applications.
- The authors developed AGENTNET TOOL, a cross-platform application that captures human computer-use demonstrations, including screen video, mouse and keyboard signals, and accessibility trees, across Windows, macOS, and Ubuntu.
- They built AGENTNET, the first large-scale and diverse computer-use task dataset with 22,625 trajectories spanning multiple operating systems and more than 200 applications and websites.
- To improve agent planning and error recovery, they propose a scalable training pipeline that integrates a new reflective long-chain-of-thought synthesis and multi-image history context encoding.
Paper links
External research summaries. These are not HDATF publications or measured product results.