Denser $\neq$ Better: Limits of On-Policy Self-Distillation for Continual Post-Training
- Published
- Source
- arXiv
- Paper number
- 555
- Field
- Machine Learning
- arXiv ID
- 2607.01763
Key points
- SDPO is effective for single-domain specialization, but it is weak in OOD generalization and continual learning.
- Compared with GRPO, SDPO causes larger changes in both parameter drift, measured by NSS and Principal Angle, and response embedding shift.
- At 5% SDPO, it exhibits an infinite repetition collapse of the boxed token, showing an artifact amplification mechanism in token-level supervision.
- Theorem 1 proves that SDPO's KL drift is always greater than or equal to that of the GRPO KL-minimal Razor policy.
- The key lesson is that on-policy data does not equal forgetting mitigation; the density of the objective function determines stability.
- Reordering data or masking tokens only partially alleviates the collapse, indicating a fundamental limitation.
Paper links
External research summaries. These are not HDATF publications or measured product results.