Mulligan: Performance-Guided Data Collection for Efficient On-Robot Learning

Published
Source
arXiv
Paper number
1167
Field
Robotics
arXiv ID
2610.05882

Key points

  • They observed that 95% of failures concentrate in 25% of the initial states, so uniform collection wastes operator time.
  • By retrying failed states and filling in unexplored ones, the robot learned faster under the same collection budget.
  • Combining a value function that also learns from failures (HiL-IDQL+Mulligan) improved final real-task success by 10-34 percentage points.
  • They validated the results on 2,550 blinded held-out episodes for higher reliability.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)