JEPA-WAM: Learning Vision-Language-Action Policies with Joint-Embedding World Modeling
- Published
- Source
- arXiv
- Paper number
- 860
- Field
- Robotics
- arXiv ID
- 2608.09381
Key points
- It proposed a latent world-action model that predicts relationships between current and future spatial structure in a pretrained V-JEPA representation space, without generating video directly.
- It combined transition prediction and continuous action generation in a single predictor so that future supervision directly refines the core structure that forms action representations.
- On LIBERO-Plus, the setting without prior robot-policy training achieved 79.2%, and the setting combined with π0.5 achieved 86.3%, each yielding the highest result.
- It also generalized to changes in visual and spatial conditions on RoboTwin 2.0 and a real dual-arm robot, showing potential for robot control without the cost of video generation.
- Because the transition objective learns shared visual structure independent of language, it may lack expressive power for tasks where different instructions produce substantially different futures from the same scene.
Paper links
External research summaries. These are not HDATF publications or measured product results.