Self-Improving Pretraining: using post-trained models to pretrain better models
- Published
- Source
- arXiv
- Paper number
- 172
- Field
- LLMs / Pretraining / Safety
- arXiv ID
- 2601.21343
Key points
- Current LLM training defers the introduction of key behaviors such as safety, factuality, quality, and complex reasoning until the post-training stage.
- Patterns learned during early pretraining lack explicit guidance for these desirable properties, which creates persistent weaknesses that are hard to fully correct later.
- Pretraining data largely lacks explicit reasoning traces, so acquiring complex reasoning ability depends mainly on expensive post-training and limits efficiency.
- The Self-Improving Pretraining framework redefines pretraining as a sequence generation task and uses stronger post-trained models as both a suffix rewriter that improves data quality and safety and a suffix judge that supplies training reward signals.
- The Think-in-the-Middle stage augments pretraining data by interleaving thinking traces, applies supervised fine-tuning, and then performs reinforcement learning to optimize the usefulness of generated thoughts for subsequent text prediction.
- It trains the policy model to generate higher-quality, safer, and more factual suffixes by using online RL algorithms such as DPO and RF-NLL with rewards supplied by the judge.
Paper links
External research summaries. These are not HDATF publications or measured product results.