Learning from the Self-future: On-policy Self-distillation for dLLMs

Published
Source
arXiv
Paper number
448
Field
LLMs / NLP
arXiv ID
2606.18195

Key points

  • It is the first on-policy self-distillation framework for dLLMs that overcomes the limitation of existing OPSD, which was designed only for autoregressive models.
  • By exploiting the non-autoregressive nature of dLLMs, it uses self-generated answers as suffix-conditioned inputs to transfer more novel knowledge.
  • It aligns dLLM's iterative denoising structure with token-level to step-level divergence supervision.
  • It achieves better reasoning performance with only about 10% of the optimization steps used by RLVR, dramatically improving sample efficiency.
  • It consistently outperforms RLVR and SFT baselines on four reasoning benchmarks.
  • It proposes a self-future-experience learning paradigm inspired by the human idea of going back to the past.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)