Matryoshka Language Model Suites
- Published
- Source
- arXiv
- Paper number
- 874
- Field
- AI / General
- arXiv ID
- 2608.09703
Key points
- We propose the Matryoshka framework, which trains submodels of different sizes end to end in a single nested structure.
- It reduces total training compute by 36% while matching the performance and validation and out-of-domain perplexity of independently trained models.
- Free distillation from the largest model to the smaller ones happens automatically at every step.
- The nested draft models improve speculative decoding throughput by 14% to 26%.
- Extensive ablations validate the importance of width and depth settings, the junction mechanism, and the distillation loss.
Paper links
External research summaries. These are not HDATF publications or measured product results.