RoboEdit: Turning Human Manipulation Videos into Scalable Robot Experience

Published
Source
arXiv
Paper number
958
Field
Robotics
arXiv ID
2608.18948

Key points

  • The key design choice is to edit human videos directly rather than generate them anew, preserving the original scene context and physical consistency.
  • Cross-embodiment adaptation modules allow a single editor to cover the appearance, kinematics, and contact dynamics of seven robot embodiments.
  • A 3D Robot-State Decoder recovers the robot hand state in every frame, providing structured supervision for downstream learning and control.
  • The automatic RoboEdit-ADC pipeline uses depth normalization and physics-based refinement to reduce artifacts such as object interpenetration, floating contacts, and temporal jitter.
  • The authors built RoboEdit-14M, a dataset of 174,000 paired videos comprising 14 million frames across seven robot embodiments, and showed that it supports training real-world manipulation policies.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)