Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Published
Source
arXiv
Paper number
033
Field
Architecture / Efficiency
arXiv ID
2502.11089

Key points

  • Traditional attention in LLMs has quadratic complexity in sequence length, making long-context processing extremely expensive during training and inference.
  • Existing sparse attention methods often fail to turn theoretical efficiency gains into real latency reductions because of inefficient memory access patterns and incompatibility with modern hardware structures such as MQA/GQA.
  • Most sparse attention approaches are inference-only and require a pretrained full-attention backbone, which can hurt performance or make end-to-end training impractical or inefficient.
  • Native Sparse Attention (NSA) introduces a hierarchical sparse attention mechanism that combines three parallel branches: compressed attention for global context, selected attention for finely important tokens, and sliding-window attention for local context.
  • A learnable gating mechanism dynamically aggregates the outputs of these branches, allowing coarse and fine information to be flexibly combined while maintaining high sparsity.
  • NSA uses dedicated Triton-based kernels designed for hardware alignment, focusing on group-centered data loading and shared KV fetching to optimize memory access and maximize GPU tensor-core utilization in both training and inference.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)