Synthetic Persona Pretraining: Alignment from Token Zero
Installing the desired assistant persona from the first token of pretraining improves constitution following, value alignment, and jailbreak robustness, and the advantage grows with pretraining budget.
TL;DR. Alignment, and the assistant identity itself, are typically introduced only after pretraining, once behavioral priors are already established, so that values are a thin overlay, rather than deeply rooted. We propose Synthetic Persona Pretraining (SPP), installing the desired assistant persona from token zero in pretraining: we annotate pretraining documents with first-person moral reflections derived from a normative value constitution, pretrain on them, and then post-train to bind the assistant identity to the pretrained persona, a phenomenon we call persona binding. Pretraining…
saved by
related reading
- Dylan Samdsam99.github.io
- Synthetic Persona Pretraining: Alignment from Token Zero — LessWronglesswrong.com
- [2601.10160] Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentarxiv.org
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- Teaching Claude Whyalignment.anthropic.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Alignment pretraining could backfire — LessWronglesswrong.com
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentalignmentpretraining.ai
- The persona selection model — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com