Learning to Move Before Learning to Do: Task-Agnostic pretraining for VLAs

Published
Source
arXiv
Paper number
559
Field
Robotics
arXiv ID
2607.02466

Key points

  • The Decomposition Hypothesis separates how to move, physical control, from what to do, semantics, and only the latter requires language.
  • An inverse-dynamics objective enables learning of a physical prior from unrelated trajectories and autonomous-driving data.
  • SIMPLER matches models trained on more than 1 million expert examples and beats behavior cloning by an absolute 10%.
  • On a real WidowX robot, it retains 25% success under camera perturbation, compared with 0% for an internet-scale baseline.
  • Partial success rises from 31.8% to 45.8%, confirming the transferability of the physical pretraining.
  • Stage 1 scale determines the performance ceiling, while extending Stage 2 has diminishing returns.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)