Trajectory-Refined Distillation
- Published
- Source
- arXiv
- Paper number
- 381
- Field
- AI / General
- arXiv ID
- 2606.08432
Key points
- It is the first to formalize the prefix-failure problem in OPD, showing that token-level improvements cannot solve it at the root.
- TRD performs distillation after refining trajectories within the on-policy support region through trajectory-level refinement.
- It can be applied to both OPD and OPSD self-distillation and yields consistent gains across the Qwen3 model family.
- On AMOBench, Qwen3-8B shows roughly 50 percent relative improvement in Pass@16, and the refined trajectories are compressed by about 9 times.
- This trajectory-level approach outperforms prior token-level loss-improvement methods such as clipping, reweighting, and top-K selection.
Paper links
External research summaries. These are not HDATF publications or measured product results.