On-Policy Self-Distillation in Diffusion Models
- Published
- Source
- arXiv
- Paper number
- 1007
- Field
- Computer Vision
- arXiv ID
- 2608.24646
Key points
- A frozen behavior policy produces trajectories and anchor points, reward gradients define bounded positive and negative targets around those anchors, and the learning policy trains toward them as detached supervision signals.
- It separates target construction from actual fitting to those targets, allowing each stage to be measured independently when performance fails to improve.
- It achieved the highest final held-out score in 19 of 20 reward-aligned conditions covering two backbones and 10 evaluators.
- Compared with DiffusionNFT, GPU time per 100 updates fell by 40% and 63%, allowing more frequent post-training to align image-generation models with human preferences under the same budget.
- However, controlled experiments showed that better target construction does not guarantee a corresponding gain from a single update: on HPSv2.1, the ranking reversed for 62.3% of 512 prompts.
Paper links
External research summaries. These are not HDATF publications or measured product results.