VIMPO: Value-Implicit Policy Optimization for LLMs

Published
Source
arXiv
Paper number
463
Field
Machine Learning
arXiv ID
2606.20008

Key points

  • It derives a closed-form expression under the optimality conditions of KL-regularized RL by expressing the value function as a policy-reference log-ratio, which removes the need for a critic.
  • It constructs a critic-free value-optimization objective with a terminal boundary condition V*(s_T) = 0 and a Monte Carlo group estimator.
  • From the same derivation, it obtains a PPO-style actor advantage, separating reward injection through the value loss from policy improvement through the actor update.
  • It shows consistent gains over GRPO on MATH-500, AIME 2024, AIME 2025, and OlympiadBench, with especially large improvements on contest-style problems.
  • It remains consistently better than GRPO even under 25 percent reward-flip noise, showing stronger robustness to reward noise.
  • It validates the method for math RLVR with a 4B base model, and shows that tuning the actor coefficient c_A and value coefficient β determines the tradeoff between learning speed and stability.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)