Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning

Published
Source
arXiv
Paper number
387
Field
Machine Learning
arXiv ID
2606.11087

Key points

  • It proposes QGF, or Q-Guided Flow, an algorithm that performs policy optimization only at test time.
  • A single Euler-step approximation uses the critic gradient of the clean action during denoising, so backpropagation through time is not needed.
  • It avoids the instability of training-time actor-critic methods and scales favorably as model size increases.
  • It achieves better performance than Best-of-N with far fewer FLOPs.
  • It shows consistent gains over existing methods on difficult goal-conditioned RL tasks.
  • When it uses a QAM critic, it extracts policies better than QAM itself.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)