Test-Time Gradient Guidance of Flow Policies in Reinforcement Learning
- Published
- Source
- arXiv
- Paper number
- 387
- Field
- Machine Learning
- arXiv ID
- 2606.11087
Key points
- It proposes QGF, or Q-Guided Flow, an algorithm that performs policy optimization only at test time.
- A single Euler-step approximation uses the critic gradient of the clean action during denoising, so backpropagation through time is not needed.
- It avoids the instability of training-time actor-critic methods and scales favorably as model size increases.
- It achieves better performance than Best-of-N with far fewer FLOPs.
- It shows consistent gains over existing methods on difficult goal-conditioned RL tasks.
- When it uses a QAM critic, it extracts policies better than QAM itself.
Paper links
External research summaries. These are not HDATF publications or measured product results.