SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning

Published
Source
arXiv
Paper number
915
Field
LLMs / NLP
arXiv ID
2608.14277

Key points

  • It connects only the tokens that occupy the same string span in both models to the teacher model's score, and uses the student model's own score for the parts that do not match.
  • The teacher model is SU-01, which can reason over more than 100K tokens. The student models are from the Qwen3, Qwen3.5, Intern-S2, GLM-4.7, and Gemma-4 families.
  • ProofBench was scored by DeepSeek-V4-Flash, and Intern-S2's score rose from 21.70 to 44.50. AnswerBench rose from 76.03 to 80.10, and AIME25 rose from 88.33 to 95.00.
  • Although it was trained only on math data, Intern-S2's HiPhO score rose from 38.6 to 41.1, and FrontierScience Research rose from 1.7 to 5.0.
  • When the teacher-imitation loss on the end token was removed and the output-distribution gap with the initial student model was constrained, the answer truncation rate dropped almost to zero. However, Gemma's AnswerBench score fell from 68.8 to 67.5, so not all capabilities improved together.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)