VIMPO: Value-Implicit Policy Optimization for LLMs
- Published
- Source
- arXiv
- Paper number
- 463
- Field
- Machine Learning
- arXiv ID
- 2606.20008
Key points
- It derives a closed-form expression under the optimality conditions of KL-regularized RL by expressing the value function as a policy-reference log-ratio, which removes the need for a critic.
- It constructs a critic-free value-optimization objective with a terminal boundary condition V*(s_T) = 0 and a Monte Carlo group estimator.
- From the same derivation, it obtains a PPO-style actor advantage, separating reward injection through the value loss from policy improvement through the actor update.
- It shows consistent gains over GRPO on MATH-500, AIME 2024, AIME 2025, and OlympiadBench, with especially large improvements on contest-style problems.
- It remains consistently better than GRPO even under 25 percent reward-flip noise, showing stronger robustness to reward noise.
- It validates the method for math RLVR with a 4B base model, and shows that tuning the actor coefficient c_A and value coefficient β determines the tradeoff between learning speed and stability.
Paper links
External research summaries. These are not HDATF publications or measured product results.