Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal

Published
Source
arXiv
Paper number
400
Field
Machine Learning
arXiv ID
2606.12360

Key points

  • Using SAE features and two-sample hypothesis tests, we identify latent concepts in preference data, including length, helpfulness, morality, and refusal.
  • We propose a data-centric framework that explains away the identified concepts to pre-remove undesirable learning signals.
  • Training on the Dolci dataset reveals reduced jailbreak robustness, meaning that safety is worse than with standard SFT.
  • We diagnose and mitigate excessive stylization and sycophancy, and the framework can amplify model personalities such as playful behavior and formality.
  • We present a unified lens for various post-training protocols, including DPO, SFT, and data filtering.
  • The approach is grounded in the theory that causally valid features change model probabilities in a predictable way.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)