MiniMax-01: Scaling Foundation Models with Lightning Attention
- Published
- Source
- arXiv
- Paper number
- 017
- Field
- LLMs / Long Context
- arXiv ID
- 2501.08313
Key points
- Existing language models are limited by short context windows of 32K to 256K tokens.
- Linear attention is theoretically promising, but it has not been successfully scaled in practice.
- It uses a hybrid architecture that combines lightning attention and softmax attention at a 7:1 ratio.
- It uses a MoE model with 32 experts and a carefully engineered parallel computation strategy.
- It follows a multi-stage training process optimized for extended context lengths.
- Linear attention can scale successfully when it is implemented and optimized properly.
Paper links
External research summaries. These are not HDATF publications or measured product results.