On Advantage Estimates for Max@K Policy Gradients
- Published
- Source
- arXiv
- Paper number
- 346
- Field
- Machine Learning
- arXiv ID
- 2606.06080
Key points
- The paper mathematically proves that the existing PKPO advantage estimator is unbiased for policy gradients but not centered, and addresses this with a Leave-Two-Out (L2O) baseline.
- It derives the canonical advantage u_i - v_i for finite-batch max@K and unifies existing pass@K and max@K estimators under a single framework.
- It is efficient with a vectorized implementation that has O(B^2) time complexity and integrates naturally into group-based RL pipelines.
- On Llama-3.2-3B-Instruct, it reduces gradient variance by 77.4 percent relative to PKPO and improves pass@256 by 5.2 percent on Qwen2.5-Math-7B.
- It shows consistent gains over strong post-training baselines for LLMs, including GRPO and Entropy-Adv.
Paper links
External research summaries. These are not HDATF publications or measured product results.