Post-Training is About States, Not Tokens: A State Distribution View of SFT, RL, and On-Policy Distillation
- Published
- Source
- arXiv
- Paper number
- 238
- Field
- Machine Learning
- arXiv ID
- 2605.22731
Key points
- SFT is not inherently destructive, but its off-policy nature becomes risky under high pressure. Developers should monitor how closely the dataset distribution aligns with the model's natural distribution.
- On-policy learning is a stabilizer. Whether through RL or distillation, sampling from the learner's own distribution helps keep updates well behaved and preserves the model's native capabilities better than off-policy methods.
- Distillation is more than imitation. By separating states from signals, on-policy distillation can let a student surpass the teacher when supervision is dense enough, using continuations rather than one-step token logits.
Paper links
External research summaries. These are not HDATF publications or measured product results.