RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation

Published
Source
arXiv
Paper number
397
Field
Machine Learning
arXiv ID
2606.11709

Key points

  • It discovers privilege-induced style drift, where privileged conditioning concentrates the learning signal on style tokens such as Therefore and Wait.
  • Contrastive hints from correct and incorrect answers cancel style components and concentrate the signal on task tokens such as mathematical content.
  • On Qwen3-8B, it achieves 90.8 percent on AMC23, 77.5 percent on AIME24, and 69.7 percent on AIME25, outperforming GRPO and OPSD variants.
  • It addresses the training instability and response-length contraction problems of prior OPSD methods, including OPSD, SDPO, SRPO, and RLSD.
  • The contrastive principle is general, so it can be plugged into existing OPSD methods and extended to cross-model distillation.
  • K-marginalization and two-path loss aggregation are key contributors to preserving performance.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)