Small-Scale Experiments: Are We There Yet?

Published
Source
arXiv
Paper number
888
Field
Machine Learning
arXiv ID
2608.11859

Key points

  • The authors experiment from about 4 million to 268 million parameters, fitting the scaling law on models from 4 million to 34 million parameters and then validating it on larger models.
  • Testing only 4 or 16 size settings failed to reveal the law, 64 settings were inaccurate, and 256 settings recovered the correct law.
  • Small models were highly sensitive to training settings, but the sensitivity decreased as the model size grew, and near the optimum the number of influential hyperparameters fell to about one.
  • The law derived at small scale extrapolated well to models about 10 times larger, but the uncertainty grew sharply at farther scales.
  • With this method, they reproduce the large-scale result that pre-normalization becomes more advantageous as the model grows, although the mapping between loss and capability holds only when the composition of the pretraining data is kept the same.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)