ActiveMimic: Egocentric Video Pretraining with Active Perception

Published
Source
arXiv
Paper number
359
Field
Robotics
arXiv ID
2606.06194

Key points

  • Camera motion in egocentric video is reinterpreted as viewpoint action rather than noise.
  • VGGT plus SAM-3D-Body reconstruct camera and wrist trajectories from a single RGB input and build a 27-dimensional unified action representation.
  • The model jointly learns active perception and manipulation with flow matching on the large-scale Ego4D dataset.
  • On four real-robot tasks, it matches or exceeds pi0 on evaluations such as Restocking at 90.1 percent and Finding at 91.7 percent.
  • The results show that active perception capability originates from egocentric pretraining rather than from robot fine-tuning.
  • Layer-level analysis confirms that camera-motion supervision promotes representation transfer from humans to robots.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)