On-Policy Self-Distillation in Diffusion Models

Published
Source
arXiv
Paper number
1007
Field
Computer Vision
arXiv ID
2608.24646

Key points

  • A frozen behavior policy produces trajectories and anchor points, reward gradients define bounded positive and negative targets around those anchors, and the learning policy trains toward them as detached supervision signals.
  • It separates target construction from actual fitting to those targets, allowing each stage to be measured independently when performance fails to improve.
  • It achieved the highest final held-out score in 19 of 20 reward-aligned conditions covering two backbones and 10 evaluators.
  • Compared with DiffusionNFT, GPU time per 100 updates fell by 40% and 63%, allowing more frequent post-training to align image-generation models with human preferences under the same budget.
  • However, controlled experiments showed that better target construction does not guarantee a corresponding gain from a single update: on HPSv2.1, the ranking reversed for 62.3% of 512 prompts.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)