Rethinking the Divergence Regularization in LLM RL
- Published
- Source
- arXiv
- Paper number
- 393
- Field
- Machine Learning
- arXiv ID
- 2606.09821
Key points
- It replaces DPPO's hard mask with an advantage-weighted l2_2 regularization term that guarantees smooth gradients.
- It preserves the geometry of the binary-TV trust region, combining divergence-based stability with SPO-style smoothness.
- Inside the boundary, it keeps the policy gradient, and beyond the boundary it provides corrective gradients that pull the policy toward the behavior policy.
- The method was cross-validated on Qwen3-4B and Qwen3-30B-A3B, a MoE model, under BF16 and FP8 settings, and it consistently outperformed GRPO, DPPO, and SPO.
- An ablation confirmed that absolute advantage weighting by |A_t| stabilizes the trust-region boundary independently of the reward scale.
- Compared with other divergence regularizers such as KL, TV, and chi2, DRPO performs better, which suggests that gradient shape matters more than the exact form of the objective.
Paper links
External research summaries. These are not HDATF publications or measured product results.