ROVE: Unlocking Human Interventions for Humanoid Manipulation via Reinforcement Learning

Published
Source
arXiv
Paper number
431
Field
Robotics
arXiv ID
2606.17011

Key points

  • It decomposes the hesitation, adaptation, and recovery stages of human remote teleoperation intervention for humanoid robots, which supports mixed-quality data modeling.
  • Optimistic Value Estimation combines TD bootstrapping and expectile regression to extract high-value actions first.
  • It adds cross-embodiment human experience videos to critic training, which strengthens long-tail failure modes.
  • Three real task intervention rounds improve success rates from 45 to 80 percent on the whiteboard task and from 56.7 to 86.7 percent on the toaster task.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)