Dense Supervision, Sparse Updates: On the Sparsity and Geometry of On-Policy Distillation
- Published
- Source
- arXiv
- Paper number
- 412
- Field
- Machine Learning
- arXiv ID
- 2606.13657
Key points
- OPD updates are small and coordinate-sparse, and they are distributed mainly across FFN layers, with consistency across ten model pairs.
- Training only the discovered sparse subnetwork is enough to recover the full OPD reasoning performance.
- SGD underperforms AdamW, which indicates that even sparse updates need adaptive scaling because the gradient scale is heterogeneous.
- The OPD delta is numerically full-rank, but its spectrum is concentrated, with the top 16 singular values accounting for about 20 to 30 percent of the energy.
- The updates avoid the principal directions of the source weight and concentrate on low-magnitude coordinates, which resembles the pattern seen in RLVR.
- Even with dense supervision, OPD looks more like sparse on-policy editing than dense rewriting.
Paper links
External research summaries. These are not HDATF publications or measured product results.