Strong Teacher Not Needed? On Distillation in LLM Pretraining

Published
Source
arXiv
Paper number
232
Field
Machine Learning
arXiv ID
2605.23857

Key points

  • Downstream accuracy refers to performance on benchmarks such as MMLU and GSM8K.
  • Do not wait for the perfect teacher. When training a larger new model, it is beneficial to use the best model you currently have as the teacher, even if that teacher is smaller or less trained than the new model's target.
  • Prioritize compatibility. A moderately strong teacher that is structurally similar to the student may be more effective than a huge, state-of-the-art model.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)