ComputerRL: Scaling End-to-End Online Reinforcement Learning for Computer Use Agents
- Published
- Source
- arXiv
- Paper number
- 075
- Field
- Computer Use Agents
- arXiv ID
- 2508.14040
Key points
- The human-computer interface is inherently designed for people, which makes robust LLM-based GUI agents difficult to build because machine interaction is more complex.
- Existing approaches such as behavior cloning suffer from scalability limits, weak generalization, and poor error recovery on diverse and complex desktop tasks.
- Applying reinforcement learning to desktop automation has been constrained by heavy compute demands, slow convergence, instability, and entropy collapse, where policies lose exploration ability during long training.
- The API-GUI interaction paradigm integrates direct GUI interaction with programmatic API calls, and it is complemented by an LLM-driven automation workflow for API development to improve efficiency and versatility.
- To collect data efficiently, the team developed a scalable and reliable Ubuntu environment with distributed multi-node clustering infrastructure that can orchestrate thousands of concurrent virtual desktop environments.
- They implemented a fully asynchronous RL framework with multi-LLM behavior cloning, step-level Group Relative Policy Optimization (GRPO) with verifiable rewards, and an Entropulse strategy for continued exploration through periodic RL-SFT alternation.
Paper links
External research summaries. These are not HDATF publications or measured product results.