InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization

Published
Source
arXiv
Paper number
572
Field
Robotics
arXiv ID
2607.04988

Key points

  • Foresight tokens are supervised only during training by the video generation model WAN2.2-5B and removed at inference time to preserve real-time speed.
  • We keep strengthening semantic ability by continuously training the VLM backbone with VQA, subtask prediction, and discrete action-token objectives.
  • It ranks first across six simulation benchmarks, including LIBERO at 98.9%, RoboTwin at 93.2%, and SimplerEnv at 80.8%.
  • It achieves 84.8% zero-shot on LIBERO-Plus and 27.7% zero-shot on DOMINO, demonstrating robustness to distribution shift and dynamic object interaction.
  • Removing foresight tokens drops DOMINO to 23.8%, confirming the value of transferring dynamic priors.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)