Trajectory-Refined Distillation

Published
Source
arXiv
Paper number
381
Field
AI / General
arXiv ID
2606.08432

Key points

  • It is the first to formalize the prefix-failure problem in OPD, showing that token-level improvements cannot solve it at the root.
  • TRD performs distillation after refining trajectories within the on-policy support region through trajectory-level refinement.
  • It can be applied to both OPD and OPSD self-distillation and yields consistent gains across the Qwen3 model family.
  • On AMOBench, Qwen3-8B shows roughly 50 percent relative improvement in Pass@16, and the refined trajectories are compressed by about 9 times.
  • This trajectory-level approach outperforms prior token-level loss-improvement methods such as clipping, reweighting, and top-K selection.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)