Rubric-Guided Self-Distillation: Post-Training Without Rubric Verifiers
- Published
- Source
- arXiv
- Paper number
- 406
- Field
- Machine Learning
- arXiv ID
- 2606.12507
Key points
- Its key idea is token-level self-distillation from a rubric-conditioned teacher to an unconditioned student.
- It uses the condition enhancement effect, where rubric-only information boosts base model performance by 30 to 45 percentage points.
- Compared with GRPO, it achieves medical gains of +6.1 versus +5.9 percentage points and science gains of +4.9 versus +4.5 percentage points, while making zero judge calls.
- On Qwen-2.5-7B, it roughly halves response length compared with GRPO and avoids verbosity drift.
- It performs on-policy student rollout distillation with clipped Jensen-Shannon divergence.
- Because GRPO with a stronger judge can outperform RGSD in some settings, the two methods are complementary.
Paper links
External research summaries. These are not HDATF publications or measured product results.