Scaling Domain Data Repetition in LLM Pretraining

Published
Source
arXiv
Paper number
919
Field
AI / General
arXiv ID
2608.14071

Key points

  • It tested four domains, code, math, wiki, and medical, using unique domain data ratios of 2.5%, 5%, and 10% of the total training set, and repeated each dataset from 1 to 7 times.
  • The Pearson correlation between the optimal number of repeats and the minimum validation loss was -0.944. It was 0.400 with model size and 0.018 with unique data ratio.
  • The optimal number of repeats for math data was about 5 to 6, and for medical data about 3 to 4. With the same total amount of domain data, math tolerated up to four repeats with little loss increase, but wiki performance degraded noticeably beyond two repeats.
  • If the total ratio of math data to general web data is fixed, validation loss on ArXiv and news data stays mostly stable even as the repeat count changes. Lowering the learning rate earlier caused performance degradation to start with fewer repeats.
  • Each training run repeated only one domain. It does not yet provide a formula for interactions when multiple domains are repeated together or for predicting the exact repeat count of a large model from small-model results.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)