Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher

Published
Source
arXiv
Paper number
1044
Field
Computer Vision
arXiv ID
2608.26872

Key points

  • It requires no teacher model at all, reaching the target in approximately 44–48 hours with a warm-start configuration, compared with DiffusionOPD's total of 97 hours (85.8h of teacher training + 11.3h of distillation), making it nearly 2 times faster.
  • K stochastic branches + a self-baseline provide dense learning signals at every step, avoiding the high-variance problem of reinforcement learning methods (such as Flow-GRPO) that use only terminal rewards.
  • Best-of-K, which follows only the good branches, was unstable during training. Pull-push over all branches + KL regularization was essential for stable convergence (an unconstrained repulsion term caused performance to collapse to 0.768).
  • It combined multiple objectives (text, composition, and aesthetics) at the reward level so that a single image could satisfy them simultaneously without gradient conflicts (a preference-score drop of 0.48 vs 1.23 with the teacher-based method).

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)