Training for the Model You Return: Improving Optimization for Iterate-Averaged Language Models
- Published
- Source
- arXiv
- Paper number
- 509
- Field
- Machine Learning
- arXiv ID
- 2606.25086
Key points
- Formulate iterate-average optimization as a continuous-time stochastic quadratic optimal control problem.
- PACE is a lightweight wrapper on top of AdamW, with extra memory equal to one copy of the model weights.
- The standard SCO convergence rate is guaranteed in the convex setting, and arbitrarily large bounded errors can be improved in the quadratic setting.
- It shows consistent gains in three fine-tuning settings on SmolLM2, Qwen3, and Gemma3, and also works for GPT-2 FineWeb pretraining.
- It is competitive with Schedule-Free and removes the need for LR decay, outperforming WSD across all token budgets with constant LR.
- Extensive ablations on pullback strength c, EMA power κ, and update frequency show its robustness.
Paper links
External research summaries. These are not HDATF publications or measured product results.