SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL
- Published
- Source
- arXiv
- Paper number
- 618
- Field
- AI / General
- arXiv ID
- 2607.11185
Key points
- VeriGen automatically synthesizes verifiable computer-use tasks through Docker interactions and feedback from multiple agents.
- It uses Frontier Sampling, which allocates executions to each task's current capability frontier, and Visual Context Segmentation, which processes only recent screens in separate segments.
- It created more than 24,000 verifiable tasks and approximately 3,000 high-quality reinforcement-learning tasks, increasing training speed by 2.83 times.
- ScaleCUA scored 68.7% on OSWorld and 54.0% on ScienceBoard and can be used for large-scale online training of open computer-use agents.
- Each execution is limited to 50 turns, and evaluation centers on the Ubuntu desktop, so generalization to very long-horizon tasks and Windows or macOS has not been verified.
Paper links
External research summaries. These are not HDATF publications or measured product results.