Learning from the Self-future: On-policy Self-distillation for dLLMs
- Published
- Source
- arXiv
- Paper number
- 448
- Field
- LLMs / NLP
- arXiv ID
- 2606.18195
Key points
- It is the first on-policy self-distillation framework for dLLMs that overcomes the limitation of existing OPSD, which was designed only for autoregressive models.
- By exploiting the non-autoregressive nature of dLLMs, it uses self-generated answers as suffix-conditioned inputs to transfer more novel knowledge.
- It aligns dLLM's iterative denoising structure with token-level to step-level divergence supervision.
- It achieves better reasoning performance with only about 10% of the optimization steps used by RLVR, dramatically improving sample efficiency.
- It consistently outperforms RLVR and SFT baselines on four reasoning benchmarks.
- It proposes a self-future-experience learning paradigm inspired by the human idea of going back to the past.
Paper links
External research summaries. These are not HDATF publications or measured product results.