LARA: Latent Action Representation Alignment for Vision-Language-Action Models
- Published
- Source
- arXiv
- Paper number
- 366
- Field
- Computer Vision
- arXiv ID
- 2606.07100
Key points
- We overcome the limitations of existing separate training pipelines by jointly optimizing LAM and VLA with a bidirectional representation-alignment loss.
- LAM is grounded in real action trajectories, which suppresses learning spurious visual variations such as background and lighting.
- VLA is regularized by LAM's forward-dynamics prediction, reducing hallucinated trajectories that are kinematically feasible but functionally wrong.
- We find that inserting alignment at DiT layer L-2 is optimal, and that alignment in deeper layers is effective.
- We achieve consistent performance gains on three simulation benchmarks and in a real robot manipulation environment.
Paper links
External research summaries. These are not HDATF publications or measured product results.