ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents

Published
Source
arXiv
Paper number
075
Field
Computer Use Agents
arXiv ID
2508.14040

Key points

  • The human-computer interface is inherently designed for people, which makes robust LLM-based GUI agents difficult to build because machine interaction is more complex.
  • Existing approaches such as behavior cloning suffer from scalability limits, weak generalization, and poor error recovery on diverse and complex desktop tasks.
  • Applying reinforcement learning to desktop automation has been constrained by heavy compute demands, slow convergence, instability, and entropy collapse, where policies lose exploration ability during long training.
  • The API-GUI interaction paradigm integrates direct GUI interaction with programmatic API calls, and it is complemented by an LLM-driven automation workflow for API development to improve efficiency and versatility.
  • To collect data efficiently, the team developed a scalable and reliable Ubuntu environment with distributed multi-node clustering infrastructure that can orchestrate thousands of concurrent virtual desktop environments.
  • They implemented a fully asynchronous RL framework with multi-LLM behavior cloning, step-level Group Relative Policy Optimization (GRPO) with verifiable rewards, and an Entropulse strategy for continued exploration through periodic RL-SFT alternation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)