CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators

Published
Source
arXiv
Paper number
1037
Field
Robotics
arXiv ID
2608.27406

Key points

  • It developed a cross-embodiment world-model framework that jointly trains a single model on videos of different robot platforms and humans.
  • It unified different action spaces through end-effector coordinates, natural language, and latent action representations.
  • Curriculum learning progressing from unlabeled video with latent actions to end-effector actions enabled zero-shot deployment.
  • In challenging environments such as DROID, it approached or surpassed single-embodiment state-of-the-art methods, with few-shot adaptation widening the gap further.
  • Planning at inference time within the world model and RL fine-tuning improved the real-world task success rates of π0.5 and MolmoAct-2.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)