Matryoshka Language Model Suites

Published
Source
arXiv
Paper number
874
Field
AI / General
arXiv ID
2608.09703

Key points

  • We propose the Matryoshka framework, which trains submodels of different sizes end to end in a single nested structure.
  • It reduces total training compute by 36% while matching the performance and validation and out-of-domain perplexity of independently trained models.
  • Free distillation from the largest model to the smaller ones happens automatically at every step.
  • The nested draft models improve speculative decoding throughput by 14% to 26%.
  • Extensive ablations validate the importance of width and depth settings, the junction mechanism, and the distillation loss.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)