Sliding-window beats linear attention
- Published
- Source
- arXiv
- Paper number
- 1046
- Field
- LLMs / NLP
- arXiv ID
- 2608.28444
Key points
- 0 additional training runs: SWA, which changes only the attention mask, matched or outperformed conversions to linear attention that require retraining on billions of tokens.
- The gap was larger in long contexts. At a 256K context length, SWA recovered 20–25% of the original performance, whereas linear attention (LoLCATs) recovered only 2.2–5%.
- On MMLU, SWA recovered 93.2% of the original performance, outperforming all major linearization methods, including LoLCATs (83.2%) and Liger-GLA (62.2%).
- SWA (window 64) had the fastest decoding speed and lowest memory use among the tested methods.
- Citations to related work also confirmed that training-free sliding windows are a strong baseline for video generation models.
Paper links
External research summaries. These are not HDATF publications or measured product results.