RLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation
- Published
- Source
- arXiv
- Paper number
- 397
- Field
- Machine Learning
- arXiv ID
- 2606.11709
Key points
- It discovers privilege-induced style drift, where privileged conditioning concentrates the learning signal on style tokens such as Therefore and Wait.
- Contrastive hints from correct and incorrect answers cancel style components and concentrate the signal on task tokens such as mathematical content.
- On Qwen3-8B, it achieves 90.8 percent on AMC23, 77.5 percent on AIME24, and 69.7 percent on AIME25, outperforming GRPO and OPSD variants.
- It addresses the training instability and response-length contraction problems of prior OPSD methods, including OPSD, SDPO, SRPO, and RLSD.
- The contrastive principle is general, so it can be plugged into existing OPSD methods and extended to cross-model distillation.
- K-marginalization and two-path loss aggregation are key contributors to preserving performance.
Paper links
External research summaries. These are not HDATF publications or measured product results.