Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning
- Published
- Source
- arXiv
- Paper number
- 583
- Field
- Machine Learning
- arXiv ID
- 2607.07508
Key points
- We replace GRPO's group-wise sampling with single-rollout to remove waiting latency and reduce off-policy drift in asynchronous RL.
- Direct double-sided IS uses rollout log-probabilities directly to remove the cost of tracking the old policy.
- skip-observation GAE skips observation tokens and computes advantage between actions in multi-turn agent trajectories.
- We update the value model more frequently than the actor and fine-tune frozen attention to stabilize single-rollout training.
- Deployed in the GLM-5.2 (750B-A40B) RL pipeline, it achieves stable training for about 1,000 steps and consistent gains over GRPO.
- In a simulated online-learning environment, it adapts quickly when reward criteria change, demonstrating robustness to nonstationary environments.
Paper links
External research summaries. These are not HDATF publications or measured product results.