Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation

Published
Source
arXiv
Paper number
412
Field
Machine Learning
arXiv ID
2606.13657

Key points

  • OPD updates are small and coordinate-sparse, and they are distributed mainly across FFN layers, with consistency across ten model pairs.
  • Training only the discovered sparse subnetwork is enough to recover the full OPD reasoning performance.
  • SGD underperforms AdamW, which indicates that even sparse updates need adaptive scaling because the gradient scale is heterogeneous.
  • The OPD delta is numerically full-rank, but its spectrum is concentrated, with the top 16 singular values accounting for about 20 to 30 percent of the energy.
  • The updates avoid the principal directions of the source weight and concentrate on low-magnitude coordinates, which resembles the pattern seen in RLVR.
  • Even with dense supervision, OPD looks more like sparse on-policy editing than dense rewriting.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)