On-policy Distillation with Verifiable Reward

Published
Source
arXiv
Paper number
1009
Field
Machine Learning
arXiv ID
2608.24696

Key points

  • It reformulates the implicit reward of sampled-token distillation in terms of trajectory correctness and applies a ReLU gate that retains only nonnegative rewards for correct trajectories and only nonpositive rewards for incorrect trajectories.
  • In a same-architecture setting with Qwen3-4B, the average across six benchmarks reached 49.1, above conventional distillation's 47.8, and its AIME24 score of 36.9 also exceeded the teacher model's 36.0.
  • A control experiment reversing the gate direction underperformed basic distillation on all six benchmarks, showing that aligning the sign with correctness is the source of the effect.
  • It can be applied directly to existing policy-gradient algorithms such as GRPO without introducing additional hyperparameters, allowing integration into existing post-training pipelines with few changes.
  • However, the student did not catch up with the teacher: in a different-architecture setting, its average of 22.8 remained far below the teacher's 30.9, and experiments were limited to mathematical-reasoning benchmarks with the Qwen3 family.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)