Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning

Published
Source
arXiv
Paper number
981
Field
Robotics
arXiv ID
2608.21204

Key points

  • By freezing the large robot policy, with billions of parameters, and updating only a Q-function of approximately 1 billion parameters, it made self-improvement cost proportional to evaluator size rather than policy size.
  • The Q-function can learn from both successful and failed rollouts, raising real-world task performance to 90%/80% where SFT retrained only on successful data had plateaued at 55%/30%.
  • At inference time, it handles 64 candidate actions sampled by the policy with a single planning step of Q-score-weighted averaging, taking 400ms and fitting within the real-time control budget.
  • Under the same online budget, it was the only method that reliably improved from failures when compared with Best-of-N, filtered SFT, IBRL, DSRL, and DAWR.
  • It also clearly stated its limitations: it still cannot solve tasks where the base policy produces no successful candidates, and it requires an episode-level success verifier.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)