Action with Visual Primitives

Published
Source
arXiv
Paper number
221
Field
Robotics
arXiv ID
2605.22183

Key points

  • Instruction and scene understanding: a pretrained VLM processes images and text to produce multimodal context tokens that capture semantic relationships between commands, such as "pick up the red chess piece," and visual scenes.
  • Spatial compositional generalization: the model was trained to move pieces in two steps, such as moving piece A to C and then C to B. At test time, it was asked to move A directly to B, a sequence it had never seen before. Baseline models failed completely, with a 0 percent success rate, while AVP retained an 83 percent success rate. Because the VLM could infer a new where-to primitive for the direct path, the action expert could execute how even though the exact trajectory was novel.
  • Cross-domain generalization: AVP was tested on unseen object categories and different backgrounds, such as switching from a chessboard to a plain white cloth.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)