QF3: Fast Flow RL with Filtered Q-Gradients
- Published
- Source
- arXiv
- Paper number
- 1170
- Field
- Robotics
- arXiv ID
- 2610.08789
Key points
- The critic's gradient is passed through a one-step prediction of the flow policy, directly improving the policy without costly differentiation through the sampler.
- Learning is stabilized by a 'filter' that cuts off action dimensions where the gradient moves too far from the replay action.
- A walking and 3-minute dance motion-tracking policy for the 29-DoF Unitree G1 humanoid was trained from scratch in simulation and transferred to real hardware without correction, a first for off-policy flow RL.
- Training is 10x faster in wall-clock time than the recent on-policy method FPO++, and comparable to the classic TD3.
- Fine-tuning a pretrained robot-arm policy made the bottle-gathering task finish episodes 22% faster while raising the success rate from 91% to 95%.
Paper links
External research summaries. These are not HDATF publications or measured product results.