Strong Teacher Not Needed? On Distillation in LLM Pretraining
- Published
- Source
- arXiv
- Paper number
- 232
- Field
- Machine Learning
- arXiv ID
- 2605.23857
Key points
- Downstream accuracy refers to performance on benchmarks such as MMLU and GSM8K.
- Do not wait for the perfect teacher. When training a larger new model, it is beneficial to use the best model you currently have as the teacher, even if that teacher is smaller or less trained than the new model's target.
- Prioritize compatibility. A moderately strong teacher that is structurally similar to the student may be more effective than a huge, state-of-the-art model.
Paper links
External research summaries. These are not HDATF publications or measured product results.