Beyond the Stability-Exploration Dilemma: Environmental Regularization for LLM Policy Optimization

Published
Source
arXiv
Paper number
990
Field
LLMs / NLP
arXiv ID
2608.23311

Key points

  • It formalized the problem that existing reinforcement learning regulates only the response distribution while allowing the question distribution to drift without control.
  • It added an input-side regularizer called Query-KL to prevent drift in the training environment while preserving freedom to explore responses.
  • It can be added to GRPO, PPO, and REINFORCE pipelines with minimal changes and no additional forward passes.
  • Replacing standard Policy-KL regularization consistently improved accuracy on 6 mathematical-reasoning benchmarks.
  • However, the paper states that ERPO is not entirely immune to collapse during long training runs, and that validation is limited to mathematical-reasoning benchmarks and Qwen-family models, leaving transfer to conversation, code generation, and multilingual settings unverified.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)