Gaze Heads: How VLMs Look at What They Describe
- Published
- Source
- arXiv
- Paper number
- 427
- Field
- Computer Vision
- arXiv ID
- 2606.14703
Key points
- About 100 of the 1,152 attention heads in a VLM, or 8.7 percent, act as gaze heads that track the image region currently being described.
- A single attention-mask intervention achieves 83.1 percent accuracy on comic panel redirection, whereas a random head fails.
- The same mechanism appears across models with 2B to 32B parameters and across VLMs such as Qwen2-VL, Ovis1.5, and InternVL3.5.
- Frozen-encoder families such as LLaVA and Bunny-3B do not show gaze heads, which suggests that joint training of the vision encoder is necessary.
Paper links
External research summaries. These are not HDATF publications or measured product results.