Post-Training is About States, Not Tokens: A State Distribution View of SFT, RL, and On-Policy Distillation

Published
Source
arXiv
Paper number
238
Field
Machine Learning
arXiv ID
2605.22731

Key points

  • SFT is not inherently destructive, but its off-policy nature becomes risky under high pressure. Developers should monitor how closely the dataset distribution aligns with the model's natural distribution.
  • On-policy learning is a stabilizer. Whether through RL or distillation, sampling from the learner's own distribution helps keep updates well behaved and preserves the model's native capabilities better than off-policy methods.
  • Distillation is more than imitation. By separating states from signals, on-policy distillation can let a student surpass the teacher when supervision is dense enough, using continuations rather than one-step token logits.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)