Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

Published
Source
arXiv
Paper number
283
Field
Machine Learning
arXiv ID
2605.29548

Key points

  • The result points to data-induced competition for resources, or neurons.
  • The paper develops a simple phenomenological argument that power-law scaling already suggests larger models can learn parts of the data distribution that smaller models cannot, even with infinite training data.
  • To test this claim and identify its cause, the authors study model scaling in a synthetic setting made of a mixture of tasks with monotonic scaling curves.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)