Variable-Width Transformers
- Published
- Source
- arXiv
- Paper number
- 445
- Field
- LLMs / NLP
- arXiv ID
- 2606.18246
Key points
- It systematically explores non-uniform width allocation by breaking the assumption that every transformer layer must have the same width.
- An X-shaped structure, wide at the ends and narrow in the middle, consistently outperforms equal-width baselines with matched parameter counts.
- Parameter-free residual resizing implements dimensional reduction and recovery without projection layers.
- With the same number of parameters, it improves perplexity by about 3 percent, reduces KV cache by about 10 percent, and reduces FLOPs by about 3 percent.
- Under a loss-matched scaling curve, it reduces FLOPs by 22 percent and extends to MoE transformers.
- It mitigates representation collapse in the middle layers and encourages qualitatively different representations at different depths.
Paper links
External research summaries. These are not HDATF publications or measured product results.