Multi-Turn On-Policy Distillation with Prefix Replay

Published
Source
arXiv
Paper number
578
Field
Machine Learning
arXiv ID
2607.04763

Key points

  • It reuses teacher RL rollouts without environment interaction, making the cost zero and requiring no separate collection as a free by-product.
  • It identifies the prefix trap as both a temporal problem, caused by error accumulation, and a distributional problem, caused by the student's occupancy shifting away from the teacher's confidence distribution.
  • Step-decay sampling gives higher weight to early steps with smaller shifts, prioritizing more reliable regions of the trajectory.
  • On math and Python tasks, when the teacher-student gap is large, ReOPD outperforms OPD, while on retrieval environments it matches OPD.
  • It is more than 4x faster per training stage than OPD, uses zero tool calls, and is easier to integrate across heterogeneous environments.
  • Prefixes generated by the teacher model itself work best, while prefixes from stronger models can actually hurt performance.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)