HumanCLAW: Can Vision-Language Models Act Through a Body?

Published
Source
arXiv
Paper number
759
Field
Computer Vision
arXiv ID
2607.27180

Key points

  • We built a framework that separates behavioral decision making in VLMs from motor control, so it evaluates reasoning ability alone.
  • We built HumanCLAW-Bench, consisting of 1,218 egocentric exploration and interaction episodes across 41 indoor scenes.
  • Even the best of nine state-of-the-art VLMs reaches only a 16.8% success rate.
  • The core failure is not perception but a lack of self-awareness, meaning the model does not know where its body is, whether it has arrived, or whether it has collided.
  • The conclusion is that current VLMs only describe the world; they do not feel the body they control.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)