Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data
- Published
- Source
- arXiv
- Paper number
- 1081
- Field
- Machine Learning
- arXiv ID
- 2609.11917
Key points
- The researchers varied data repetition rates, domain mixtures, and expert counts and sizes across models with 80 million to 1 billion active parameters.
- At the 80-million-active-parameter scale, mixture-of-experts models degraded from four data repetitions onward and lost their advantage over dense models at high repetition rates.
- Expert-routing paths freezing early in training and experts over-specializing on repeated data were observed, offering clues for understanding the overfitting.
- Dropout and output masking mitigated the damage from repeated data, suggesting that overfitting countermeasures deserve review alongside architecture choices when data is limited.
- None of the tested mitigations fully reproduced training on unique data, so repeated-data training cannot be treated as a complete substitute for new data.
Paper links
External research summaries. These are not HDATF publications or measured product results.