Why do LLMs attend to the first token?
- Published
- Source
- arXiv
- Paper number
- 050
- Field
- Interpretability
- arXiv ID
- 2504.02732
Key points
- The persistent, unexplained phenomenon in which large language models disproportionately attend to the first token, called an attention sink, lacked a functional explanation.
- Information degradation in deep Transformers, including rank collapse, representation collapse, and over-squashing, causes loss of individual token information, especially in long sequences.
- Prior literature did not provide a fundamental understanding of why attention sinks occur or what functional purpose they serve.
- This paper theoretically connects attention sinks to excessive mixing and representation collapse, and derives a new over-squashing bound that quantifies how sinks regulate information flow.
- We performed extensive empirical validation on state-of-the-art LLMs Gemma 7B and the LLaMa 3.1 family through perturbation analysis, close inspection of attention patterns, and targeted ablation studies.
- Controlled pretraining experiments demonstrated a causal relationship between context length, model scale, and the formation and strength of attention sinks.
Paper links
External research summaries. These are not HDATF publications or measured product results.