LARA: Latent Action Representation Alignment for Vision-Language-Action Models

Published
Source
arXiv
Paper number
366
Field
Computer Vision
arXiv ID
2606.07100

Key points

  • We overcome the limitations of existing separate training pipelines by jointly optimizing LAM and VLA with a bidirectional representation-alignment loss.
  • LAM is grounded in real action trajectories, which suppresses learning spurious visual variations such as background and lighting.
  • VLA is regularized by LAM's forward-dynamics prediction, reducing hallucinated trajectories that are kinematically feasible but functionally wrong.
  • We find that inserting alignment at DiT layer L-2 is optimal, and that alignment in deeper layers is effective.
  • We achieve consistent performance gains on three simulation benchmarks and in a real robot manipulation environment.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)