Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
- Published
- Source
- arXiv
- Paper number
- 033
- Field
- Architecture / Efficiency
- arXiv ID
- 2502.11089
Key points
- Traditional attention in LLMs has quadratic complexity in sequence length, making long-context processing extremely expensive during training and inference.
- Existing sparse attention methods often fail to turn theoretical efficiency gains into real latency reductions because of inefficient memory access patterns and incompatibility with modern hardware structures such as MQA/GQA.
- Most sparse attention approaches are inference-only and require a pretrained full-attention backbone, which can hurt performance or make end-to-end training impractical or inefficient.
- Native Sparse Attention (NSA) introduces a hierarchical sparse attention mechanism that combines three parallel branches: compressed attention for global context, selected attention for finely important tokens, and sliding-window attention for local context.
- A learnable gating mechanism dynamically aggregates the outputs of these branches, allowing coarse and fine information to be flexibly combined while maintaining high sparsity.
- NSA uses dedicated Triton-based kernels designed for hardware alignment, focusing on group-centered data loading and shared KV fetching to optimize memory access and maximize GPU tensor-core utilization in both training and inference.
Paper links
External research summaries. These are not HDATF publications or measured product results.