Weak-to-Strong On-Policy Distillation
- Published
- Source
- arXiv
- Paper number
- 761
- Field
- Machine Learning
- arXiv ID
- 2607.26246
Key points
- It proposes a new distillation method that separates the capability direction from the logit difference of a weak model pair, positive and negative, and adds it to a strong student.
- It proposes three settings: RL before and after model pairs, model pairs of different sizes, and answer versus hint pairs.
- Across four math and three coding benchmarks, it achieves relative improvements over existing OPD of 11.4% in math and 12.0% in coding.
- Even though it learns only from weak models, it outperforms the actual weak teacher models themselves.
- It found that the RL-based pair and the size-based pair provide complementary signals, strengthening the reasoning skeleton and the solution procedure, respectively.
Paper links
External research summaries. These are not HDATF publications or measured product results.