MiniMax-01: Scaling Foundation Models with Lightning Attention

Published
Source
arXiv
Paper number
017
Field
LLMs / Long Context
arXiv ID
2501.08313

Key points

  • Existing language models are limited by short context windows of 32K to 256K tokens.
  • Linear attention is theoretically promising, but it has not been successfully scaled in practice.
  • It uses a hybrid architecture that combines lightning attention and softmax attention at a 7:1 ratio.
  • It uses a MoE model with 32 experts and a carefully engineered parallel computation strategy.
  • It follows a multi-stage training process optimized for extended context lengths.
  • Linear attention can scale successfully when it is implemented and optimized properly.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)