Beyond Imitation: Self-Improving Robot Policies via Off-Policy Q-Planning
- Published
- Source
- arXiv
- Paper number
- 981
- Field
- Robotics
- arXiv ID
- 2608.21204
Key points
- By freezing the large robot policy, with billions of parameters, and updating only a Q-function of approximately 1 billion parameters, it made self-improvement cost proportional to evaluator size rather than policy size.
- The Q-function can learn from both successful and failed rollouts, raising real-world task performance to 90%/80% where SFT retrained only on successful data had plateaued at 55%/30%.
- At inference time, it handles 64 candidate actions sampled by the policy with a single planning step of Q-score-weighted averaging, taking 400ms and fitting within the real-time control budget.
- Under the same online budget, it was the only method that reliably improved from failures when compared with Best-of-N, filtered SFT, IBRL, DSRL, and DAWR.
- It also clearly stated its limitations: it still cannot solve tasks where the base policy produces no successful candidates, and it requires an episode-level success verifier.
Paper links
External research summaries. These are not HDATF publications or measured product results.