World Tokens: Enhancing Embodied Policies with Training-Time World Modeling
- Published
- Source
- arXiv
- Paper number
- 868
- Field
- Computer Vision
- arXiv ID
- 2608.09730
Key points
- It proposes the World Tokens architecture, which uses a video world model only during training and removes it at deployment.
- The World Adapter converts VLM features into 256 tokens and conditions both video prediction and action generation on them.
- The action expert cannot see the full VLM sequence directly, which prevents it from bypassing the world-model supervision.
- With a 2B model, it reaches 98.2 percent on LIBERO, 71.5 percent on SIMPLER-WidowX, and 82.1 percent on SIMPLER-GoogleRobot.
- On the real R1 Pro robot, it improves success from 59.4 percent to 76.0 percent compared with the matching baseline.
Paper links
External research summaries. These are not HDATF publications or measured product results.