Variable-Width Transformers

Published
Source
arXiv
Paper number
445
Field
LLMs / NLP
arXiv ID
2606.18246

Key points

  • It systematically explores non-uniform width allocation by breaking the assumption that every transformer layer must have the same width.
  • An X-shaped structure, wide at the ends and narrow in the middle, consistently outperforms equal-width baselines with matched parameter counts.
  • Parameter-free residual resizing implements dimensional reduction and recovery without projection layers.
  • With the same number of parameters, it improves perplexity by about 3 percent, reduces KV cache by about 10 percent, and reduces FLOPs by about 3 percent.
  • Under a loss-matched scaling curve, it reduces FLOPs by 22 percent and extends to MoE transformers.
  • It mitigates representation collapse in the middle layers and encourages qualitatively different representations at different depths.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)