Rethinking the Divergence Regularization in LLM RL

Published
Source
arXiv
Paper number
393
Field
Machine Learning
arXiv ID
2606.09821

Key points

  • It replaces DPPO's hard mask with an advantage-weighted l2_2 regularization term that guarantees smooth gradients.
  • It preserves the geometry of the binary-TV trust region, combining divergence-based stability with SPO-style smoothness.
  • Inside the boundary, it keeps the policy gradient, and beyond the boundary it provides corrective gradients that pull the policy toward the behavior policy.
  • The method was cross-validated on Qwen3-4B and Qwen3-30B-A3B, a MoE model, under BF16 and FP8 settings, and it consistently outperformed GRPO, DPPO, and SPO.
  • An ablation confirmed that absolute advantage weighting by |A_t| stabilizes the trust-region boundary independently of the reward scale.
  • Compared with other divergence regularizers such as KL, TV, and chi2, DRPO performs better, which suggests that gradient shape matters more than the exact form of the objective.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)