Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models

Published
Source
arXiv
Paper number
1003
Field
Robotics
arXiv ID
2608.23478

Key points

  • It goes beyond behavior cloning by distilling the local purpose, or intent, of actions as an explicit supervision signal.
  • Teacher inference and target generation occur only once offline, and all teacher-side modules are removed at deployment, so adding it to an existing VLA increases runtime overhead only by the additional decoder representations.
  • It improved GR00T-N1.7 on SimplerEnv-Bridge from 64.3% to 84.7% (+20.4 percentage points) and on RoboCasa from 64.1% to 70.3%.
  • Average performance on real-robot tasks rose from 62.0% to 68.7%, with gains of up to +12 percentage points on long-horizon tasks.
  • However, the meaning of intent is constrained by what the teacher VLM knows: with a teacher lacking embodied experience, targets lost almost all directionality, and performance fell below even the baseline that learns only by imitating actions.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)