SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning
- Published
- Source
- arXiv
- Paper number
- 915
- Field
- LLMs / NLP
- arXiv ID
- 2608.14277
Key points
- It connects only the tokens that occupy the same string span in both models to the teacher model's score, and uses the student model's own score for the parts that do not match.
- The teacher model is SU-01, which can reason over more than 100K tokens. The student models are from the Qwen3, Qwen3.5, Intern-S2, GLM-4.7, and Gemma-4 families.
- ProofBench was scored by DeepSeek-V4-Flash, and Intern-S2's score rose from 21.70 to 44.50. AnswerBench rose from 76.03 to 80.10, and AIME25 rose from 88.33 to 95.00.
- Although it was trained only on math data, Intern-S2's HiPhO score rose from 38.6 to 41.1, and FrontierScience Research rose from 1.7 to 5.0.
- When the teacher-imitation loss on the end token was removed and the output-distribution gap with the initial student model was constrained, the answer truncation rate dropped almost to zero. However, Gemma's AnswerBench score fell from 68.8 to 67.5, so not all capabilities improved together.
Paper links
External research summaries. These are not HDATF publications or measured product results.