Sliding-window beats linear attention

Published
Source
arXiv
Paper number
1046
Field
LLMs / NLP
arXiv ID
2608.28444

Key points

  • 0 additional training runs: SWA, which changes only the attention mask, matched or outperformed conversions to linear attention that require retraining on billions of tokens.
  • The gap was larger in long contexts. At a 256K context length, SWA recovered 20–25% of the original performance, whereas linear attention (LoLCATs) recovered only 2.2–5%.
  • On MMLU, SWA recovered 93.2% of the original performance, outperforming all major linearization methods, including LoLCATs (83.2%) and Liger-GLA (62.2%).
  • SWA (window 64) had the fastest decoding speed and lowest memory use among the tested methods.
  • Citations to related work also confirmed that training-free sliding windows are a strong baseline for video generation models.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)