Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

Published
Source
arXiv
Paper number
893
Field
LLMs / NLP
arXiv ID
2608.12149

Key points

  • The first pattern is a large spike in values immediately before a full attention layer. The second pattern is those values staying high across several linear attention layers between spikes.
  • The same patterns are observed in five linear-attention architectures, six mixture ratios, five input domains, and public models ranging from 1.2B to 397B parameters.
  • In a controlled 340M and 1.3B parameter 12-to-1 mixture, alignment between the layer right before full attention and the spike position was 99.4% to 100% overall.
  • In the 340M control model, both patterns appeared from the first measurement point, after 1B training tokens. The full-attention output gate reduced magnitude but did not eliminate the layer-wise position pattern.
  • The authors explain that fast cancellation makes spikes and slow cancellation makes plateaus, but they leave the cause of the cancellation timing and its actual computational role for future work.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)