GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Published
Source
arXiv
Paper number
102
Field
RL / Training
arXiv ID
2601.05242

Key points

  • Aligning large language models with diverse human preferences, such as safety, consistency, and efficiency, requires multi-reward reinforcement learning.
  • The commonly used Group Relative Policy Optimization, or GRPO, algorithm suffers from reward-signal collapse when applied to the sum of multiple rewards, meaning that different reward combinations can produce the same advantage value.
  • This collapse reduces the resolution of the learning signal, leading to suboptimal convergence, poorer training stability, and potential early training failure.
  • GDPO preserves fine-grained distinctions by decoupling normalization of individual reward components and performing group-wise normalization separately for each reward.
  • The individually normalized advantages are then summed and passed through an additional batch-level normalization step to ensure numerical stability during policy updates.
  • To manage priority across rewards, GDPO explores reward weighting and proposes reward-function conditioning, which gives easier rewards only when harder goals have been achieved.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)