AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding

Published
Source
arXiv
Paper number
360
Field
Robotics
arXiv ID
2606.06155

Key points

  • It introduces a structured affordance intermediate representation to close the gap between the semantic space of VLMs and the 3D physical action space.
  • It models the task progressively through Which2Act, a visual latent prediction stage, Where2Act, a 2D affordance map stage, and How2Act, a 3D geometry stage.
  • Its MoT architecture consists of three expert modules for understanding, affordance generation, and action.
  • It addresses label sparsity with a three-stage progressive training scheme and an automatic affordance annotation pipeline.
  • On real robots, it achieves 88.3 percent on basic tasks and 82.9 percent on complex tasks, which is 38.1 percentage points better than Pi0 on complex tasks.
  • The authors hypothesize that affordances serve as a semantic anchor that preserves the visual-language capabilities of the VLM backbone.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)