HumanCLAW: Can Vision-Language Models Act Through a Body?
- Published
- Source
- arXiv
- Paper number
- 759
- Field
- Computer Vision
- arXiv ID
- 2607.27180
Key points
- We built a framework that separates behavioral decision making in VLMs from motor control, so it evaluates reasoning ability alone.
- We built HumanCLAW-Bench, consisting of 1,218 egocentric exploration and interaction episodes across 41 indoor scenes.
- Even the best of nine state-of-the-art VLMs reaches only a 16.8% success rate.
- The core failure is not perception but a lack of self-awareness, meaning the model does not know where its body is, whether it has arrived, or whether it has collided.
- The conclusion is that current VLMs only describe the world; they do not feel the body they control.
Paper links
External research summaries. These are not HDATF publications or measured product results.