Rubric-Guided Self-Distillation: Post-Training Without Rubric Verifiers

Published
Source
arXiv
Paper number
406
Field
Machine Learning
arXiv ID
2606.12507

Key points

  • Its key idea is token-level self-distillation from a rubric-conditioned teacher to an unconditioned student.
  • It uses the condition enhancement effect, where rubric-only information boosts base model performance by 30 to 45 percentage points.
  • Compared with GRPO, it achieves medical gains of +6.1 versus +5.9 percentage points and science gains of +4.9 versus +4.5 percentage points, while making zero judge calls.
  • On Qwen-2.5-7B, it roughly halves response length compared with GRPO and avoids verbosity drift.
  • It performs on-policy student rollout distillation with clipped Jensen-Shannon divergence.
  • Because GRPO with a stronger judge can outperform RGSD in some settings, the two methods are complementary.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)