Self-OPD: On-Policy Distillation for Flow Matching Models without Teacher
- Published
- Source
- arXiv
- Paper number
- 1044
- Field
- Computer Vision
- arXiv ID
- 2608.26872
Key points
- It requires no teacher model at all, reaching the target in approximately 44–48 hours with a warm-start configuration, compared with DiffusionOPD's total of 97 hours (85.8h of teacher training + 11.3h of distillation), making it nearly 2 times faster.
- K stochastic branches + a self-baseline provide dense learning signals at every step, avoiding the high-variance problem of reinforcement learning methods (such as Flow-GRPO) that use only terminal rewards.
- Best-of-K, which follows only the good branches, was unstable during training. Pull-push over all branches + KL regularization was essential for stable convergence (an unconstrained repulsion term caused performance to collapse to 0.768).
- It combined multiple objectives (text, composition, and aesthetics) at the reward level so that a single image could satisfy them simultaneously without gradient conflicts (a preference-score drop of 0.48 vs 1.23 with the teacher-based method).
Paper links
External research summaries. These are not HDATF publications or measured product results.