SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL

Published
Source
arXiv
Paper number
618
Field
AI / General
arXiv ID
2607.11185

Key points

  • VeriGen automatically synthesizes verifiable computer-use tasks through Docker interactions and feedback from multiple agents.
  • It uses Frontier Sampling, which allocates executions to each task's current capability frontier, and Visual Context Segmentation, which processes only recent screens in separate segments.
  • It created more than 24,000 verifiable tasks and approximately 3,000 high-quality reinforcement-learning tasks, increasing training speed by 2.83 times.
  • ScaleCUA scored 68.7% on OSWorld and 54.0% on ScienceBoard and can be used for large-scale online training of open computer-use agents.
  • Each execution is limited to 50 turns, and evaluation centers on the Ubuntu desktop, so generalization to very long-horizon tasks and Windows or macOS has not been verified.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)