MiniMax Sparse Attention

Published
Source
arXiv
Paper number
403
Field
AI / General
arXiv ID
2606.13392

Key points

  • It proposes block sparse attention based on GQA, where a lightweight index branch independently selects Top-k blocks for each GQA group.
  • From-scratch training on 3T tokens in a 109B MoE model demonstrates downstream performance equivalent to GQA.
  • At 1M context length, it reduces attention FLOPs per token by 28.4x and speeds up prefill by 14.2x and decoding by 7.6x on H800.
  • It optimizes GPU tensor-core utilization with an exp-free TopK kernel and KV-outer sparse attention.
  • The index branch is trained with a KL alignment loss and separated from the main branch with stop-gradient.
  • It is publicly released as the commercial multimodal model MiniMax-M3 on HuggingFace.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)