Long Context Pre-Training with Lighthouse Attention

Published
Source
arXiv
Paper number
180
Field
LLMs / Architecture / Long Context
arXiv ID
2605.06554

Key points

  • The quadratic time and memory complexity of scaled dot-product attention, Theta(N^2), makes it computationally infeasible to train causal Transformers on extremely long sequences such as 128K or 1M-plus tokens.
  • Existing sparse-attention methods for pretraining are often asymmetric, pooling only keys and values, or structurally entangled, which prevents reuse of highly optimized dense attention kernels such as FlashAttention.
  • It remains a major challenge to ensure that a model pretrained with sparse attention can function effectively as a full dense-attention model at inference time.
  • The paper introduces a four-stage pipeline that wraps the attention kernel so standard FlashAttention can operate on dynamically compressed subsequences: pyramid construction, scoring and selection, gather-sequence attention, and scatter-restore reconstruction.
  • It applies symmetric average pooling to queries, keys, and values across multiple pyramid levels, and combines a parameter-free projection-norm scorer with a non-differentiable top-k selection step for quasi-quadratic processing.
  • It uses a two-stage training scheme: first train with Lighthouse Attention for efficiency, then run a short resume stage with full dense SDPA at inference time to restore and optimize the model's full attention capacity.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)