Action with Visual Primitives
- Published
- Source
- arXiv
- Paper number
- 221
- Field
- Robotics
- arXiv ID
- 2605.22183
Key points
- Instruction and scene understanding: a pretrained VLM processes images and text to produce multimodal context tokens that capture semantic relationships between commands, such as "pick up the red chess piece," and visual scenes.
- Spatial compositional generalization: the model was trained to move pieces in two steps, such as moving piece A to C and then C to B. At test time, it was asked to move A directly to B, a sequence it had never seen before. Baseline models failed completely, with a 0 percent success rate, while AVP retained an 83 percent success rate. Because the VLM could infer a new where-to primitive for the direct path, the action expert could execute how even though the exact trajectory was novel.
- Cross-domain generalization: AVP was tested on unseen object categories and different backgrounds, such as switching from a chessboard to a plain white cloth.
Paper links
External research summaries. These are not HDATF publications or measured product results.