InSight: Self-Guided Skill Acquisition via Steerable VLAs

Published
Source
arXiv
Paper number
482
Field
Robotics
arXiv ID
2606.24884

Key points

  • It automatically splits human demonstrations into VLM planning segments and end-effector trajectories, making the VLA controllable at the primitive-action level.
  • The VLM identifies missing primitive actions in a new task and builds a flywheel that tries them through low-level control, collects success data, and retrains.
  • It achieves 92% and 96% end-to-end success rates on twisting and following without human demonstrations, compared with 32% and 16% for CaP-X.
  • It reaches 80% success on long-horizon tasks chained from 14 primitive actions, such as twisting followed by following.
  • Adding new primitive actions does not reduce existing pick-and-place skill, confirming that forgetting is not a problem.
  • In a Mars-robot scenario, it learns a sweeping behavior autonomously from only mopping demonstrations and succeeds on all 5 evaluations.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)