A First-Principles Derivation of LLM Policy Optimization: From Expected Reward to GRPO and Its Structural Extensions
- Published
- Source
- arXiv
- Paper number
- 435
- Field
- AI / General
- arXiv ID
- 2606.16733
Key points
- It unifies all policy gradient methods into two axes, trajectory side and reward side, for J(theta) = E[R(tau)].
- It explains the evolution from REINFORCE to PPO to GRPO to post-GRPO variants through the failure modes and interventions at each stage.
- It extends the same two-axis view to agentic RL and hybrid GRPO-OPD.
- It identifies coupled failures that cannot be solved by changing only one axis and points to directions for next-generation algorithm design.
Paper links
External research summaries. These are not HDATF publications or measured product results.