World Tokens: Enhancing Embodied Policies with Training-Time World Modeling

Published
Source
arXiv
Paper number
868
Field
Computer Vision
arXiv ID
2608.09730

Key points

  • It proposes the World Tokens architecture, which uses a video world model only during training and removes it at deployment.
  • The World Adapter converts VLM features into 256 tokens and conditions both video prediction and action generation on them.
  • The action expert cannot see the full VLM sequence directly, which prevents it from bypassing the world-model supervision.
  • With a 2B model, it reaches 98.2 percent on LIBERO, 71.5 percent on SIMPLER-WidowX, and 82.1 percent on SIMPLER-GoogleRobot.
  • On the real R1 Pro robot, it improves success from 59.4 percent to 76.0 percent compared with the matching baseline.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)