Attention to Mamba: A Recipe for Cross-Architecture Distillation

Published
Source
arXiv
Paper number
152
Field
LLMs / Architecture / Distillation
arXiv ID
2604.14191

Key points

  • Transformer models are powerful, but they incur quadratic compute cost as sequence length grows, which limits efficiency on long sequences and inference.
  • Alternative linear architectures, including state-space models such as Mamba, offer efficiency but often underperform Transformers on downstream tasks at scale.
  • Previous attempts at cross-architecture distillation from Transformers to Mamba-like structures struggled to preserve performance and often required hybrid attention-SSM solutions.
  • The paper develops a principled two-stage distillation recipe that transfers first from Softmax Attention to Linear Attention (Hedgehog), and then from Linear Attention to adapted Mamba (HedgeMamba).
  • The Mamba SSM mixer parameters are initialized using the feature maps learned in the Linear Attention stage, and additional Mamba components are set to act as identity operators.
  • The full HedgeMamba structure is then fine-tuned with standard cross-entropy loss on the answers, activating all Mamba components.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)