ActiveMimic: Egocentric Video Pretraining with Active Perception
- Published
- Source
- arXiv
- Paper number
- 359
- Field
- Robotics
- arXiv ID
- 2606.06194
Key points
- Camera motion in egocentric video is reinterpreted as viewpoint action rather than noise.
- VGGT plus SAM-3D-Body reconstruct camera and wrist trajectories from a single RGB input and build a 27-dimensional unified action representation.
- The model jointly learns active perception and manipulation with flow matching on the large-scale Ego4D dataset.
- On four real-robot tasks, it matches or exceeds pi0 on evaluations such as Restocking at 90.1 percent and Finding at 91.7 percent.
- The results show that active perception capability originates from egocentric pretraining rather than from robot fine-tuning.
- Layer-level analysis confirms that camera-motion supervision promotes representation transfer from humans to robots.
Paper links
External research summaries. These are not HDATF publications or measured product results.