Weak-to-Strong On-Policy Distillation

Published
Source
arXiv
Paper number
761
Field
Machine Learning
arXiv ID
2607.26246

Key points

  • It proposes a new distillation method that separates the capability direction from the logit difference of a weak model pair, positive and negative, and adds it to a strong student.
  • It proposes three settings: RL before and after model pairs, model pairs of different sizes, and answer versus hint pairs.
  • Across four math and three coding benchmarks, it achieves relative improvements over existing OPD of 11.4% in math and 12.0% in coding.
  • Even though it learns only from weak models, it outperforms the actual weak teacher models themselves.
  • It found that the RL-based pair and the size-based pair provide complementary signals, strengthening the reasoning skeleton and the solution procedure, respectively.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)