HumanScale: Egocentric Human Video Can Outperform Real-Robot Data for Embodied Pretraining
- Published
- Source
- arXiv
- Paper number
- 461
- Field
- Computer Vision
- arXiv ID
- 2606.20521
Key points
- Egocentric pretraining shows a clear scaling law, with validation loss decreasing monotonically as more data is added.
- At the same 5,000-hour scale, egocentric data reduces validation loss by 24% versus robot data, improves in-distribution success from 40% to 92.5%, and improves OOD success from 0% to 90%.
- The comparison is based on a World-Action Model with a MoT architecture, using a framework that combines video generation and action prediction.
- Egocentric video has lower normalized jerk, meaning smoother motion, and a lower idle fraction, which makes the data more efficient.
- In real-robot experiments on an AgiBot bi-manipulator, the egocentric pretraining model maintained a 90% success rate even on OOD objects, while the baseline collapsed to 0%.
- Robot data shows almost no improvement in unseen-task generalization even when scaled up, demonstrating the fundamental limitations of real robot data.
Paper links
External research summaries. These are not HDATF publications or measured product results.