PC Layer: Polynomial Weight Preconditioning for Improving LLM Pre-Training

Published
Source
arXiv
Paper number
336
Field
Machine Learning
arXiv ID
2606.06470

Key points

  • A low-order polynomial preconditioner stabilizes the singular value spectrum of the weight matrices.
  • The trained weights are merged back into the original architecture, so there is no inference overhead.
  • On Llama-1B, it improves performance over the standard transformer for both AdamW and Muon.
  • Theoretically, it guarantees geometric convergence of gradient descent when layerwise singular values are uniformly bounded.
  • It can be applied to existing LLM training pipelines without extra parameters or architectural changes.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)