GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
- Published
- Source
- arXiv
- Paper number
- 102
- Field
- RL / Training
- arXiv ID
- 2601.05242
Key points
- Aligning large language models with diverse human preferences, such as safety, consistency, and efficiency, requires multi-reward reinforcement learning.
- The commonly used Group Relative Policy Optimization, or GRPO, algorithm suffers from reward-signal collapse when applied to the sum of multiple rewards, meaning that different reward combinations can produce the same advantage value.
- This collapse reduces the resolution of the learning signal, leading to suboptimal convergence, poorer training stability, and potential early training failure.
- GDPO preserves fine-grained distinctions by decoupling normalization of individual reward components and performing group-wise normalization separately for each reward.
- The individually normalized advantages are then summed and passed through an additional batch-level normalization step to ensure numerical stability during policy updates.
- To manage priority across rewards, GDPO explores reward weighting and proposes reward-function conditioning, which gives easier rewards only when harder goals have been achieved.
Paper links
External research summaries. These are not HDATF publications or measured product results.