Why do LLMs attend to the first token?

Published
Source
arXiv
Paper number
050
Field
Interpretability
arXiv ID
2504.02732

Key points

  • The persistent, unexplained phenomenon in which large language models disproportionately attend to the first token, called an attention sink, lacked a functional explanation.
  • Information degradation in deep Transformers, including rank collapse, representation collapse, and over-squashing, causes loss of individual token information, especially in long sequences.
  • Prior literature did not provide a fundamental understanding of why attention sinks occur or what functional purpose they serve.
  • This paper theoretically connects attention sinks to excessive mixing and representation collapse, and derives a new over-squashing bound that quantifies how sinks regulate information flow.
  • We performed extensive empirical validation on state-of-the-art LLMs Gemma 7B and the LLaMa 3.1 family through perturbation analysis, close inspection of attention patterns, and targeted ablation studies.
  • Controlled pretraining experiments demonstrated a causal relationship between context length, model scale, and the formation and strength of attention sinks.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)