Flux-OPD: On-Policy Distillation with Evolving Contexts

Published
Source
arXiv
Paper number
783
Field
Machine Learning
arXiv ID
2607.28022

Key points

  • It proposes Flux-OPD, which uses evolving contexts as supervision signals during training in open-ended domains.
  • Through reverse KL decomposition, it shows that the student is distilled by a geometric mean teacher and defines a conflict term.
  • The context correction strategy does not use an unstable context teacher as the direct target, and instead injects only the difference signal into a stable reference teacher.
  • It uses the conflict term as a weighting signal, correcting strongly when contexts agree and weakly when they conflict.
  • It achieves 19.63 to 20.61 on HealthBench and 81.48 to 82.26 on VBench for 1.3B, with consistent gains across three independent runs.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)