Synthetic Persona Pretraining: Alignment from Token Zero

Published
Source
arXiv
Paper number
898
Field
Machine Learning
arXiv ID
2608.13482

Key points

  • It proposes persona binding, which attaches value reflections to pretraining documents to install a persona early and then binds it to the assistant identity during post-training.
  • Injecting from the first token improves constitutional compliance and leads to more aligned choices even on moral dilemmas not covered during training.
  • If the same reflection is injected only at the end of pretraining, the effect is weak and value priorities do not change.
  • The effect grows as the pretraining budget increases, from 1.7B/100B to 3B/500B.
  • Late injection alone was enough for jailbreak defense, but deep value alignment required early injection.

Paper links

External research summaries. These are not HDATF publications or measured product results.

Read original (opens in a new tab)