AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning

Published
Source
arXiv
Paper number
083
Field
Agents / RL
arXiv ID
2509.08755

Key points

  • There is no integrated, interactive reinforcement learning (RL) framework that can train LLM agents from scratch for complex, multi-turn, long-horizon tasks.
  • Existing methods for developing LLM agents, such as prompting and supervised fine-tuning (SFT), often rely on proprietary models or human-curated data, limiting scalability, adaptability, and intrinsic self-improvement.
  • Traditional RL applications to LLMs are mostly limited to single-turn tasks or face instability and inefficiency when extended to multi-turn agentic decision making.
  • AgentGym-RL is a modular open-source framework that separates the environment, agent, and training modules to support diverse multi-turn long-horizon tasks.
  • ScalingInter-RL is a new progressive interactive scaling strategy that adaptively increases the maximum number of turns during RL training to balance exploration and exploitation while improving optimization stability and efficiency.
  • The framework integrates a full RL pipeline that supports mainstream online RL algorithms such as PPO and GRPO, with engineering optimizations for scalability and reliability across a variety of realistic environments.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)