Attention to Mamba: A Recipe for Cross-Architecture Distillation
- Published
- Source
- arXiv
- Paper number
- 152
- Field
- LLMs / Architecture / Distillation
- arXiv ID
- 2604.14191
Key points
- Transformer models are powerful, but they incur quadratic compute cost as sequence length grows, which limits efficiency on long sequences and inference.
- Alternative linear architectures, including state-space models such as Mamba, offer efficiency but often underperform Transformers on downstream tasks at scale.
- Previous attempts at cross-architecture distillation from Transformers to Mamba-like structures struggled to preserve performance and often required hybrid attention-SSM solutions.
- The paper develops a principled two-stage distillation recipe that transfers first from Softmax Attention to Linear Attention (Hedgehog), and then from Linear Attention to adapted Mamba (HedgeMamba).
- The Mamba SSM mixer parameters are initialized using the feature maps learned in the Linear Attention stage, and additional Mamba components are set to act as identity operators.
- The full HedgeMamba structure is then fine-tuned with standard cross-entropy loss on the answers, activating all Mamba components.
Paper links
External research summaries. These are not HDATF publications or measured product results.