flâneur

Synthetic Persona Pretraining: Alignment from Token Zero

modelraising.ai · 3,901 words · saved by 1 readers

Installing the desired assistant persona from the first token of pretraining improves constitution following, value alignment, and jailbreak robustness, and the advantage grows with pretraining budget.

TL;DR. Alignment, and the assistant identity itself, are typically introduced only after pretraining, once behavioral priors are already established, so that values are a thin overlay, rather than deeply rooted. We propose Synthetic Persona Pretraining (SPP), installing the desired assistant persona from token zero in pretraining: we annotate pretraining documents with first-person moral reflections derived from a normative value constitution, pretrain on them, and then post-train to bind the assistant identity to the pretrained persona, a phenomenon we call persona binding. Pretraining…

saved by

related reading