Self-Improving Pretraining: using post-trained models to pretrain better models

Published
Source
arXiv
Paper number
172
Field
LLMs / Pretraining / Safety
arXiv ID
2601.21343

Key points

  • Current LLM training defers the introduction of key behaviors such as safety, factuality, quality, and complex reasoning until the post-training stage.
  • Patterns learned during early pretraining lack explicit guidance for these desirable properties, which creates persistent weaknesses that are hard to fully correct later.
  • Pretraining data largely lacks explicit reasoning traces, so acquiring complex reasoning ability depends mainly on expensive post-training and limits efficiency.
  • The Self-Improving Pretraining framework redefines pretraining as a sequence generation task and uses stronger post-trained models as both a suffix rewriter that improves data quality and safety and a suffix judge that supplies training reward signals.
  • The Think-in-the-Middle stage augments pretraining data by interleaving thinking traces, applies supervised fine-tuning, and then performs reinforcement learning to optimize the usefulness of generated thoughts for subsequent text prediction.
  • It trains the policy model to generate higher-quality, safer, and more factual suffixes by using online RL algorithms such as DPO and RF-NLL with rewards supplied by the judge.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)