InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization
- Published
- Source
- arXiv
- Paper number
- 572
- Field
- Robotics
- arXiv ID
- 2607.04988
Key points
- Foresight tokens are supervised only during training by the video generation model WAN2.2-5B and removed at inference time to preserve real-time speed.
- We keep strengthening semantic ability by continuously training the VLM backbone with VQA, subtask prediction, and discrete action-token objectives.
- It ranks first across six simulation benchmarks, including LIBERO at 98.9%, RoboTwin at 93.2%, and SimplerEnv at 80.8%.
- It achieves 84.8% zero-shot on LIBERO-Plus and 27.7% zero-shot on DOMINO, demonstrating robustness to distribution shift and dynamic object interaction.
- Removing foresight tokens drops DOMINO to 23.8%, confirming the value of transferring dynamic priors.
Paper links
External research summaries. These are not HDATF publications or measured product results.