On the Geometry of On-Policy Distillation

Published
Source
arXiv
Paper number
363
Field
Machine Learning
arXiv ID
2606.07082

Key points

  • OPD lies in a relaxed off-principal regime between SFT, which aligns principal components, and RLVR, which localizes non-principal components.
  • The authors discover subspace locking, in which updates lock into a narrow low-dimensional channel soon after training begins.
  • Under a rank-16 projection constraint, OPD preserves performance while SFT degrades sharply, which shows that the locked subspace is functionally sufficient.
  • Token sparsification and off-policy rollouts preserve the rank trajectory, but objective mixing changes it, which confirms that objective composition is the key control axis.
  • Controlled experiments based on Qwen3-8B use multidimensional diagnostic metrics such as update sparsity, spectral drift, and principal-component rotation.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)