Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal
- Published
- Source
- arXiv
- Paper number
- 400
- Field
- Machine Learning
- arXiv ID
- 2606.12360
Key points
- Using SAE features and two-sample hypothesis tests, we identify latent concepts in preference data, including length, helpfulness, morality, and refusal.
- We propose a data-centric framework that explains away the identified concepts to pre-remove undesirable learning signals.
- Training on the Dolci dataset reveals reduced jailbreak robustness, meaning that safety is worse than with standard SFT.
- We diagnose and mitigate excessive stylization and sycophancy, and the framework can amplify model personalities such as playful behavior and formality.
- We present a unified lens for various post-training protocols, including DPO, SFT, and data filtering.
- The approach is grounded in the theory that causally valid features change model probabilities in a predictable way.
Paper links
External research summaries. These are not HDATF publications or measured product results.