Small-Scale Experiments: Are We There Yet?
- Published
- Source
- arXiv
- Paper number
- 888
- Field
- Machine Learning
- arXiv ID
- 2608.11859
Key points
- The authors experiment from about 4 million to 268 million parameters, fitting the scaling law on models from 4 million to 34 million parameters and then validating it on larger models.
- Testing only 4 or 16 size settings failed to reveal the law, 64 settings were inaccurate, and 256 settings recovered the correct law.
- Small models were highly sensitive to training settings, but the sensitivity decreased as the model size grew, and near the optimum the number of influential hyperparameters fell to about one.
- The law derived at small scale extrapolated well to models about 10 times larger, but the uncertainty grew sharply at farther scales.
- With this method, they reproduce the large-scale result that pre-normalization becomes more advantageous as the model grows, although the mapping between loss and capability holds only when the composition of the pretraining data is kept the same.
Paper links
External research summaries. These are not HDATF publications or measured product results.