Do as I Do: Dexterous Manipulation Data from Everyday Human Videos
- Published
- Source
- arXiv
- Paper number
- 449
- Field
- Robotics
- arXiv ID
- 2606.19333
Key points
- It proposes a guided diffusion method that repurposes the SAM 3D generation model as a video object tracker, achieving tracking performance 67% better than FoundationPose.
- It achieves state-of-the-art reconstruction accuracy on the DexYCB and HOI4D benchmarks, with F-5 of 0.71 and 0.72 and CD of 0.66 and 0.49.
- Three novel components, warmup, perturbation, and transition reward, improve retargeting success from 25% to 71%.
- It also shows effectiveness on clean data, improving OakInk2 MoCap performance from 72% to 81%.
- Analysis of the 100DOH dataset shows that only about 4% of online human videos can be directly used for dexterous manipulation learning.
- It validates real execution on a two-arm system with a 22-DoF Sharpa Wave hand and UR3e arm across 10 tasks, including stirring, pouring, and wiping.
Paper links
External research summaries. These are not HDATF publications or measured product results.