LA4VLA: Learning to Act without Seeing via Language-Action Pretraining

Published
Source
arXiv
Paper number
508
Field
Robotics
arXiv ID
2606.27295

Key points

  • It provides a systematic analysis of vision-language asymmetry in VLA training at both the data level and the input level.
  • It automatically builds 33K LA episodes from existing demonstrations without collecting extra data, in LA4-33K.
  • LA-only pretraining consistently outperforms VLA-only pretraining at the same scale.
  • Mixed LA-VLA pretraining reaches 87.53% on MetaWorld, 96.28% on LIBERO, and 83.3% on real robots.
  • Under visual noise, the LA-pretrained policy achieves a success rate that is 25 percentage points higher than VLA-only.
  • t-SNE analysis of the LA-pretrained policy shows clearly separated representations by command direction.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)