Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human Data

Published
Source
arXiv
Paper number
793
Field
Robotics
arXiv ID
2608.02580

Key points

  • The pipeline creates virtual fingertips from 21 hand keypoints using a weighted index and middle finger combination of 0.7 and 0.3, extracts the gripper center and opening width, and removes temporal noise with Savitzky-Golay filtering and SLERP. For videos without hand-pose annotations, it estimates pose with WiLoR and DynHaMR, and Qwen3.5 splits long videos into sub-tasks.
  • From about 1,940 hours of raw video, it produces 18,561 hours of data for 15 robot morphologies, which is about a 9.6x expansion.
  • In pretraining mixture experiments, the 1:1 mix ranks first in five of seven settings, reaching RoboTwin Clean 68.1% plus 5.9, Randomized 53.5% plus 2.6, and task-semantic axis 54.1% plus 7.9, while the 1:3 mix brings almost no gain.
  • By factor, lighting adds 7.6 points, robot color adds 6.4, and camera offset adds 5.9, which means visual robustness gives the largest gain. Unseen objects rise from 29.3% to 40.0% with the 3:1 mix, paraphrasing reaches 68.5%, and Franka stays below 7%, which shows transfer to robots with very different kinematics failed.
  • In the ablation that uses only egocentric data and no robot data, raw egocentric video scores 28.1%, the pipeline raises it to 31.7%, increasing the number of morphologies from 1 to 15 raises it to 33.5%, and adding the raw egocentric data raises it further to 37.3%.
  • On five long-horizon tasks with a real ARX ACone robot, the best result came from converting 20 teleoperation demonstrations into about 7-minute egocentric play videos per scene and mixing them, which improved block insertion by 14 points and screw insertion by 13 points.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)