One Policy, Many Embodiments: Unified Camera-Centric Action Geometry Pre-training for Heterogeneous Embodied Manipulation

Published
Source
arXiv
Paper number
1022
Field
Robotics
arXiv ID
2608.26058

Key points

  • It aligned robot, simulation, and human-hand data in a single camera-centric action space for unified pretraining.
  • It directly incorporated 2,340 hours of human demonstrations into training as a 'human embodiment,' without retargeting or video synthesis.
  • A single checkpoint achieved 98.3% on LIBERO, 88.7%/89.2% on RoboTwin Easy/Hard, and 82.0% zero-shot on LIBERO-Plus.
  • On a real robot, it achieved 60% for picking up bread, 90% for opening a drawer, and 75% for stacking bowls, outperforming pi0.5 (20/85/65%) under the same conditions.
  • It also disclosed limitations: camera-calibration and depth-estimation errors propagate directly, and cross-embodiment transfer remains difficult (35% from ALOHA to ARX).

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)