Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization
- Published
- Source
- arXiv
- Paper number
- 990
- Field
- LLMs / NLP
- arXiv ID
- 2608.23311
Key points
- It formalized the problem that existing reinforcement learning regulates only the response distribution while allowing the question distribution to drift without control.
- It added an input-side regularizer called Query-KL to prevent drift in the training environment while preserving freedom to explore responses.
- It can be added to GRPO, PPO, and REINFORCE pipelines with minimal changes and no additional forward passes.
- Replacing standard Policy-KL regularization consistently improved accuracy on 6 mathematical-reasoning benchmarks.
- However, the paper states that ERPO is not entirely immune to collapse during long training runs, and that validation is limited to mathematical-reasoning benchmarks and Qwen-family models, leaving transfer to conversation, code generation, and multilingual settings unverified.
Paper links
External research summaries. These are not HDATF publications or measured product results.