Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data

Published
Source
arXiv
Paper number
1081
Field
Machine Learning
arXiv ID
2609.11917

Key points

  • The researchers varied data repetition rates, domain mixtures, and expert counts and sizes across models with 80 million to 1 billion active parameters.
  • At the 80-million-active-parameter scale, mixture-of-experts models degraded from four data repetitions onward and lost their advantage over dense models at high repetition rates.
  • Expert-routing paths freezing early in training and experts over-specializing on repeated data were observed, offering clues for understanding the overfitting.
  • Dropout and output masking mitigated the damage from repeated data, suggesting that overfitting countermeasures deserve review alongside architecture choices when data is limited.
  • None of the tested mitigations fully reproduced training on unique data, so repeated-data training cannot be treated as a complete substitute for new data.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)