JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling

Published
Source
arXiv
Paper number
860
Field
Robotics
arXiv ID
2608.09381

Key points

  • It proposed a latent world-action model that predicts relationships between current and future spatial structure in a pretrained V-JEPA representation space, without generating video directly.
  • It combined transition prediction and continuous action generation in a single predictor so that future supervision directly refines the core structure that forms action representations.
  • On LIBERO-Plus, the setting without prior robot-policy training achieved 79.2%, and the setting combined with π0.5 achieved 86.3%, each yielding the highest result.
  • It also generalized to changes in visual and spatial conditions on RoboTwin 2.0 and a real dual-arm robot, showing potential for robot control without the cost of video generation.
  • Because the transition objective learns shared visual structure independent of language, it may lack expressive power for tasks where different instructions produce substantially different futures from the same scene.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)