Self-Policy Distillation via Capability-Selective Subspace Projection

Published
Source
arXiv
Paper number
224
Field
LLMs / NLP
arXiv ID
2605.22675

Key points

  • Independence: It does not require an external teacher or reward model.
  • Generalization: It works across heterogeneous domains and transfers between tasks.
  • Selectivity: It targets the underlying capability signal rather than superficial patterns.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)