DAPD: Dual-Anchored Policy Distillation

Published
Source
arXiv
Paper number
818
Field
AI / General
arXiv ID
2608.01735

Key points

  • It diagnoses the root cause of privilege illusion as information asymmetry between the teacher and the student.
  • It uses the self-conditioned distribution as a bridge to align reference and rollout behaviors along two information-matched paths.
  • It supervises both directions, from reference to rollout and from rollout to reference, to reduce dependence on privileged information while preserving accuracy.
  • It reduces wrong claims by 45% compared with OPSD while consistently improving reasoning performance.
  • The improvement persists as the model scales from 1.7B to 32B.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)