OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation
- Published
- Source
- arXiv
- Paper number
- 349
- Field
- Machine Learning
- arXiv ID
- 2606.06096
Key points
- We propose an unbiased gradient estimator for order-statistic objective functions, and by changing only the rank weights, it can express many distributional objectives such as VaR, CVaR, trimmed mean, and best-of-K.
- It is a plug-and-play design that can be applied to existing likelihood-ratio and reparameterization updates in one line of code by reformulating them as reward transformations.
- Top-M@K shows consistent gains over the existing Max@K, which is the special case max@K with m=1, on both pass@1 and pass@256.
- Combining a correct-answer reward for Top-M and a length penalty for Bottom-M shortens responses far more effectively than the simple scalarization used by GRPO.
- Choosing K, the target size, and m, the number of top ranks, controls the exploration-exploitation tradeoff, with larger K improving pass@256 while reducing pass@1.
- It has effectively negligible overhead, with an O(N log N) sorting cost.
Paper links
External research summaries. These are not HDATF publications or measured product results.