On-policy Distillation with Verifiable Reward
- Published
- Source
- arXiv
- Paper number
- 1009
- Field
- Machine Learning
- arXiv ID
- 2608.24696
Key points
- It reformulates the implicit reward of sampled-token distillation in terms of trajectory correctness and applies a ReLU gate that retains only nonnegative rewards for correct trajectories and only nonpositive rewards for incorrect trajectories.
- In a same-architecture setting with Qwen3-4B, the average across six benchmarks reached 49.1, above conventional distillation's 47.8, and its AIME24 score of 36.9 also exceeded the teacher model's 36.0.
- A control experiment reversing the gate direction underperformed basic distillation on all six benchmarks, showing that aligning the sign with correctness is the source of the effect.
- It can be applied directly to existing policy-gradient algorithms such as GRPO without introducing additional hyperparameters, allowing integration into existing post-training pipelines with few changes.
- However, the student did not catch up with the teacher: in a different-architecture setting, its average of 22.8 remained far below the teacher's 30.9, and experiments were limited to mathematical-reasoning benchmarks with the Qwen3 family.
Paper links
External research summaries. These are not HDATF publications or measured product results.