HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining

Published
Source
arXiv
Paper number
461
Field
Computer Vision
arXiv ID
2606.20521

Key points

  • Egocentric pretraining shows a clear scaling law, with validation loss decreasing monotonically as more data is added.
  • At the same 5,000-hour scale, egocentric data reduces validation loss by 24% versus robot data, improves in-distribution success from 40% to 92.5%, and improves OOD success from 0% to 90%.
  • The comparison is based on a World-Action Model with a MoT architecture, using a framework that combines video generation and action prediction.
  • Egocentric video has lower normalized jerk, meaning smoother motion, and a lower idle fraction, which makes the data more efficient.
  • In real-robot experiments on an AgiBot bi-manipulator, the egocentric pretraining model maintained a 90% success rate even on OOD objects, while the baseline collapsed to 0%.
  • Robot data shows almost no improvement in unseen-task generalization even when scaled up, demonstrating the fundamental limitations of real robot data.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)