Denser $\neq$ Better: Limits of On-Policy Self-Distillation for Continual Post-Training

Published
Source
arXiv
Paper number
555
Field
Machine Learning
arXiv ID
2607.01763

Key points

  • SDPO is effective for single-domain specialization, but it is weak in OOD generalization and continual learning.
  • Compared with GRPO, SDPO causes larger changes in both parameter drift, measured by NSS and Principal Angle, and response embedding shift.
  • At 5% SDPO, it exhibits an infinite repetition collapse of the boxed token, showing an artifact amplification mechanism in token-level supervision.
  • Theorem 1 proves that SDPO's KL drift is always greater than or equal to that of the GRPO KL-minimal Razor policy.
  • The key lesson is that on-policy data does not equal forgetting mitigation; the density of the objective function determines stability.
  • Reordering data or masking tokens only partially alleviates the collapse, indicating a fundamental limitation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)