Distilled Reinforcement Learning for LLM Post-training
- Published
- Source
- arXiv
- Paper number
- 676
- Field
- Machine Learning
- arXiv ID
- 2607.17247
Key points
- Reinforcement learning gives one reward per response, so it is hard to know which intermediate reasoning step to praise, and it mostly reinforces what the student can already do rather than adding new knowledge.
- On-policy distillation forces the model to follow the teacher distribution through KL divergence, creating a dilemma in which there is little to learn if the teacher is too similar and too much noise if it is from a different family.
- Distilled RL does not add a separate imitation objective; instead, it redistributes the learning signal of RL token by token using the reverse importance ratio between the teacher policy and the old student policy.
- It consists of three parts: reverse importance sampling with ratio clipping, negative-sample reset that restores the weight of negative-advantage responses to 1, and response-level geometric normalization that sets the geometric mean of ratios within a response to 1.
- Without negative-sample reset, mean pass@1 drops by 8.81 points on Qwen3-4B and 6.39 points on DSQW-1.5B, making it the most important of the three parts.
- The gains are especially large when the teacher and student belong to different families. In that setting, on-policy distillation remains below RL, while Distilled RL beats all baselines.
Paper links
External research summaries. These are not HDATF publications or measured product results.