CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators
- Published
- Source
- arXiv
- Paper number
- 1037
- Field
- Robotics
- arXiv ID
- 2608.27406
Key points
- It developed a cross-embodiment world-model framework that jointly trains a single model on videos of different robot platforms and humans.
- It unified different action spaces through end-effector coordinates, natural language, and latent action representations.
- Curriculum learning progressing from unlabeled video with latent actions to end-effector actions enabled zero-shot deployment.
- In challenging environments such as DROID, it approached or surpassed single-embodiment state-of-the-art methods, with few-shot adaptation widening the gap further.
- Planning at inference time within the world model and RL fine-tuning improved the real-world task success rates of π0.5 and MolmoAct-2.
Paper links
External research summaries. These are not HDATF publications or measured product results.