Gaze Heads: How VLMs Look at What They Describe

Published
Source
arXiv
Paper number
427
Field
Computer Vision
arXiv ID
2606.14703

Key points

  • About 100 of the 1,152 attention heads in a VLM, or 8.7 percent, act as gaze heads that track the image region currently being described.
  • A single attention-mask intervention achieves 83.1 percent accuracy on comic panel redirection, whereas a random head fails.
  • The same mechanism appears across models with 2B to 32B parameters and across VLMs such as Qwen2-VL, Ovis1.5, and InternVL3.5.
  • Frozen-encoder families such as LLaVA and Bunny-3B do not show gaze heads, which suggests that joint training of the vision encoder is necessary.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)