AgentGym-RL: Training LLM Agents for Long-Horizon Decision Making through Multi-Turn Reinforcement Learning
- Published
- Source
- arXiv
- Paper number
- 083
- Field
- Agents / RL
- arXiv ID
- 2509.08755
Key points
- There is no integrated, interactive reinforcement learning (RL) framework that can train LLM agents from scratch for complex, multi-turn, long-horizon tasks.
- Existing methods for developing LLM agents, such as prompting and supervised fine-tuning (SFT), often rely on proprietary models or human-curated data, limiting scalability, adaptability, and intrinsic self-improvement.
- Traditional RL applications to LLMs are mostly limited to single-turn tasks or face instability and inefficiency when extended to multi-turn agentic decision making.
- AgentGym-RL is a modular open-source framework that separates the environment, agent, and training modules to support diverse multi-turn long-horizon tasks.
- ScalingInter-RL is a new progressive interactive scaling strategy that adaptively increases the maximum number of turns during RL training to balance exploration and exploitation while improving optimization stability and efficiency.
- The framework integrates a full RL pipeline that supports mainstream online RL algorithms such as PPO and GRPO, with engineering optimizations for scalability and reliability across a variety of realistic environments.
Paper links
External research summaries. These are not HDATF publications or measured product results.