Training for the Model You Return: Improving Optimization for Iterate-Averaged Language Models

Published
Source
arXiv
Paper number
509
Field
Machine Learning
arXiv ID
2606.25086

Key points

  • Formulate iterate-average optimization as a continuous-time stochastic quadratic optimal control problem.
  • PACE is a lightweight wrapper on top of AdamW, with extra memory equal to one copy of the model weights.
  • The standard SCO convergence rate is guaranteed in the convex setting, and arbitrarily large bounded errors can be improved in the quadratic setting.
  • It shows consistent gains in three fine-tuning settings on SmolLM2, Qwen3, and Gemma3, and also works for GPT-2 FineWeb pretraining.
  • It is competitive with Schedule-Free and removes the need for LR decay, outperforming WSD across all token budgets with constant LR.
  • Extensive ablations on pullback strength c, EMA power κ, and update frequency show its robustness.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)