Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention
- Published
- Source
- arXiv
- Paper number
- 283
- Field
- Machine Learning
- arXiv ID
- 2605.29548
Key points
- The result points to data-induced competition for resources, or neurons.
- The paper develops a simple phenomenological argument that power-law scaling already suggests larger models can learn parts of the data distribution that smaller models cannot, even with infinite training data.
- To test this claim and identify its cause, the authors study model scaling in a synthetic setting made of a mixture of tasks with monotonic scaling curves.
Paper links
External research summaries. These are not HDATF publications or measured product results.