Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning

Published
Source
arXiv
Paper number
583
Field
Machine Learning
arXiv ID
2607.07508

Key points

  • We replace GRPO's group-wise sampling with single-rollout to remove waiting latency and reduce off-policy drift in asynchronous RL.
  • Direct double-sided IS uses rollout log-probabilities directly to remove the cost of tracking the old policy.
  • skip-observation GAE skips observation tokens and computes advantage between actions in multi-turn agent trajectories.
  • We update the value model more frequently than the actor and fine-tune frozen attention to stabilize single-rollout training.
  • Deployed in the GLM-5.2 (750B-A40B) RL pipeline, it achieves stable training for about 1,000 steps and consistent gains over GRPO.
  • In a simulated online-learning environment, it adapts quickly when reward criteria change, demonstrating robustness to nonstationary environments.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)