MiniMax Sparse Attention
- Published
- Source
- arXiv
- Paper number
- 403
- Field
- AI / General
- arXiv ID
- 2606.13392
Key points
- It proposes block sparse attention based on GQA, where a lightweight index branch independently selects Top-k blocks for each GQA group.
- From-scratch training on 3T tokens in a 109B MoE model demonstrates downstream performance equivalent to GQA.
- At 1M context length, it reduces attention FLOPs per token by 28.4x and speeds up prefill by 14.2x and decoding by 7.6x on H800.
- It optimizes GPU tensor-core utilization with an exp-free TopK kernel and KV-outer sparse attention.
- The index branch is trained with a KL alignment loss and separated from the main branch with stop-gradient.
- It is publicly released as the commercial multimodal model MiniMax-M3 on HuggingFace.
Paper links
External research summaries. These are not HDATF publications or measured product results.