On the Geometry of On-Policy Distillation
- Published
- Source
- arXiv
- Paper number
- 363
- Field
- Machine Learning
- arXiv ID
- 2606.07082
Key points
- OPD lies in a relaxed off-principal regime between SFT, which aligns principal components, and RLVR, which localizes non-principal components.
- The authors discover subspace locking, in which updates lock into a narrow low-dimensional channel soon after training begins.
- Under a rank-16 projection constraint, OPD preserves performance while SFT degrades sharply, which shows that the locked subspace is functionally sufficient.
- Token sparsification and off-policy rollouts preserve the rank trajectory, but objective mixing changes it, which confirms that objective composition is the key control axis.
- Controlled experiments based on Qwen3-8B use multidimensional diagnostic metrics such as update sparsity, spectral drift, and principal-component rotation.
Paper links
External research summaries. These are not HDATF publications or measured product results.