Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus
- Published
- Source
- arXiv
- Paper number
- 893
- Field
- LLMs / NLP
- arXiv ID
- 2608.12149
Key points
- The first pattern is a large spike in values immediately before a full attention layer. The second pattern is those values staying high across several linear attention layers between spikes.
- The same patterns are observed in five linear-attention architectures, six mixture ratios, five input domains, and public models ranging from 1.2B to 397B parameters.
- In a controlled 340M and 1.3B parameter 12-to-1 mixture, alignment between the layer right before full attention and the spike position was 99.4% to 100% overall.
- In the 340M control model, both patterns appeared from the first measurement point, after 1B training tokens. The full-attention output gate reduced magnitude but did not eliminate the layer-wise position pattern.
- The authors explain that fast cancellation makes spikes and slow cancellation makes plateaus, but they leave the cause of the cancellation timing and its actual computational role for future work.
Paper links
External research summaries. These are not HDATF publications or measured product results.