LLMs as Noisy Channels: A Shannon Perspective on Model Capacity and Scaling Laws
- Published
- Source
- arXiv
- Paper number
- 228
- Field
- Machine Learning
- arXiv ID
- 2605.23901
Key points
- Irreducible noise, e, corresponds to the constant term in Chinchilla-style laws and represents a fundamental structural limit or system entropy that cannot be eliminated by scaling.
- Strategic scaling: the presence of a U-shaped curve suggests that there is an optimal sweet spot for scaling N and D at a given noise level. Blind scaling is not always the answer, and developers should consider the signal-to-noise ratio of their training setup.
- Noise awareness: exponent analysis shows that data noise often grows faster than the learning signal, with delta greater than beta. This underscores the critical importance of data curation and noise mitigation. If your training data is noisy, simply adding more of it can eventually backfire.
Paper links
External research summaries. These are not HDATF publications or measured product results.