Distilled Reinforcement Learning for LLM Post-training

Published
Source
arXiv
Paper number
676
Field
Machine Learning
arXiv ID
2607.17247

Key points

  • Reinforcement learning gives one reward per response, so it is hard to know which intermediate reasoning step to praise, and it mostly reinforces what the student can already do rather than adding new knowledge.
  • On-policy distillation forces the model to follow the teacher distribution through KL divergence, creating a dilemma in which there is little to learn if the teacher is too similar and too much noise if it is from a different family.
  • Distilled RL does not add a separate imitation objective; instead, it redistributes the learning signal of RL token by token using the reverse importance ratio between the teacher policy and the old student policy.
  • It consists of three parts: reverse importance sampling with ratio clipping, negative-sample reset that restores the weight of negative-advantage responses to 1, and response-level geometric normalization that sets the geometric mean of ratios within a response to 1.
  • Without negative-sample reset, mean pass@1 drops by 8.81 points on Qwen3-4B and 6.39 points on DSQW-1.5B, making it the most important of the three parts.
  • The gains are especially large when the teacher and student belong to different families. In that setting, on-policy distillation remains below RL, while Distilled RL beats all baselines.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)