Tapered Language Models
- Published
- Source
- arXiv
- Paper number
- 473
- Field
- Machine Learning
- arXiv ID
- 2606.23670
Key points
- It applies a cosine decay to layerwise MLP width, such as 1.5x at the front and 0.5x at the back, to concentrate capacity early and improve perplexity by up to 1.84 points over uniform allocation.
- It shows consistent gains across four architectures, Transformer, Gated Attention, Hope-attention, and Titans, and three model sizes, 440M, 760M, and 1.3B.
- It improves downstream benchmark performance without increasing the parameter count or FLOPs.
- The mechanism analysis shows that deeper MLP outputs align more strongly with the residual stream, which means they reinforce existing information more than they update it.
- It is a general design lever that is not tied to one architecture and can also be applied to other foundation models such as vision transformers and diffusion models.
Paper links
External research summaries. These are not HDATF publications or measured product results.