You Only Index Once: Cross-Layer Sparse Attention with Shared Routing
- Published
- Source
- arXiv
- Paper number
- 339
- Field
- LLMs / NLP
- arXiv ID
- 2606.06467
Key points
- It proposes CLSA, which shares routing indices across layers in a KV-sharing architecture.
- A single indexer performs top-k selection once and reuses the result across all layers.
- It amortizes routing cost while preserving the accuracy of token-sparse attention.
- It improves all major inference bottlenecks at once, including prefilling, KV cache, and decoding.
- It achieves a 7.6x decoding speedup and a 17.1x overall throughput gain at 128K context length.
Paper links
External research summaries. These are not HDATF publications or measured product results.