AffordanceVLA: A Vision-Language-Action Model Empowering Action Generation through Affordance-Aware Understanding
- Published
- Source
- arXiv
- Paper number
- 360
- Field
- Robotics
- arXiv ID
- 2606.06155
Key points
- It introduces a structured affordance intermediate representation to close the gap between the semantic space of VLMs and the 3D physical action space.
- It models the task progressively through Which2Act, a visual latent prediction stage, Where2Act, a 2D affordance map stage, and How2Act, a 3D geometry stage.
- Its MoT architecture consists of three expert modules for understanding, affordance generation, and action.
- It addresses label sparsity with a three-stage progressive training scheme and an automatic affordance annotation pipeline.
- On real robots, it achieves 88.3 percent on basic tasks and 82.9 percent on complex tasks, which is 38.1 percentage points better than Pi0 on complex tasks.
- The authors hypothesize that affordances serve as a semantic anchor that preserves the visual-language capabilities of the VLM backbone.
Paper links
External research summaries. These are not HDATF publications or measured product results.