Do as I Do: Dexterous Manipulation Data from Everyday Human Videos

Published
Source
arXiv
Paper number
449
Field
Robotics
arXiv ID
2606.19333

Key points

  • It proposes a guided diffusion method that repurposes the SAM 3D generation model as a video object tracker, achieving tracking performance 67% better than FoundationPose.
  • It achieves state-of-the-art reconstruction accuracy on the DexYCB and HOI4D benchmarks, with F-5 of 0.71 and 0.72 and CD of 0.66 and 0.49.
  • Three novel components, warmup, perturbation, and transition reward, improve retargeting success from 25% to 71%.
  • It also shows effectiveness on clean data, improving OakInk2 MoCap performance from 72% to 81%.
  • Analysis of the 100DOH dataset shows that only about 4% of online human videos can be directly used for dexterous manipulation learning.
  • It validates real execution on a two-arm system with a 22-DoF Sharpa Wave hand and UR3e arm across 10 tasks, including stirring, pouring, and wiping.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)