You Only Index Once: Cross-Layer Sparse Attention with Shared Routing

Published
Source
arXiv
Paper number
339
Field
LLMs / NLP
arXiv ID
2606.06467

Key points

  • It proposes CLSA, which shares routing indices across layers in a KV-sharing architecture.
  • A single indexer performs top-k selection once and reuses the result across all layers.
  • It amortizes routing cost while preserving the accuracy of token-sparse attention.
  • It improves all major inference bottlenecks at once, including prefilling, KV cache, and decoding.
  • It achieves a 7.6x decoding speedup and a 17.1x overall throughput gain at 128K context length.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)