DAPD: Dual-Anchored Policy Distillation
- Published
- Source
- arXiv
- Paper number
- 818
- Field
- AI / General
- arXiv ID
- 2608.01735
Key points
- It diagnoses the root cause of privilege illusion as information asymmetry between the teacher and the student.
- It uses the self-conditioned distribution as a bridge to align reference and rollout behaviors along two information-matched paths.
- It supervises both directions, from reference to rollout and from rollout to reference, to reduce dependence on privileged information while preserving accuracy.
- It reduces wrong claims by 45% compared with OPSD while consistently improving reasoning performance.
- The improvement persists as the model scales from 1.7B to 32B.
Paper links
External research summaries. These are not HDATF publications or measured product results.