Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models
- Published
- Source
- arXiv
- Paper number
- 1003
- Field
- Robotics
- arXiv ID
- 2608.23478
Key points
- It goes beyond behavior cloning by distilling the local purpose, or intent, of actions as an explicit supervision signal.
- Teacher inference and target generation occur only once offline, and all teacher-side modules are removed at deployment, so adding it to an existing VLA increases runtime overhead only by the additional decoder representations.
- It improved GR00T-N1.7 on SimplerEnv-Bridge from 64.3% to 84.7% (+20.4 percentage points) and on RoboCasa from 64.1% to 70.3%.
- Average performance on real-robot tasks rose from 62.0% to 68.7%, with gains of up to +12 percentage points on long-horizon tasks.
- However, the meaning of intent is constrained by what the teacher VLM knows: with a teacher lacking embodied experience, targets lost almost all directionality, and performance fell below even the baseline that learns only by imitating actions.
Paper links
External research summaries. These are not HDATF publications or measured product results.