Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWrong
LLMs pretrained on data about misaligned AIs themselves become less aligned. Luckily, pretraining LLMs with synthetic data about good AIs helps them become more aligned. These alignment priors persist through post-training, providing alignment-in-depth. We recommend labs pretrain for alignment, just as they do for capabilities. Website: alignmentpretraining.ai Us: geodesicresearch.org | x.com/geodesresearch Note: We are currently garnering feedback here before submitting to ICML. Any suggestions here or on our Google Doc are welcome! We will be releasing a revision on arXiv in the coming days. Folks who leave feedback will be added to the Acknowledgment section. Thank you! We pretrained a suite of 6.9B-parameter LLMs, varying only the content related to AI systems, and evaluated them for misalignment. When filtering the vast majority of the content related to AI, we see significant decreases in misalignment rates. The opposite was also true - synthetic positive AI data led to self-fulf
x Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWrong Alignment Pretraining Self Fulfilling/Refuting Prophecies LLM Personas Aligned AI Role-Model Fiction AI Frontpage 2025 Top Fifty: 23 % 201 Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment by Cam , Puria , Kyle O’Brien , David Africa , Samuel Ratnam , andyk 21st Dec 2025 AI Alignment Forum 10 min read 25 201 Ω 59 TL;DR LLMs pretrained on data about misaligned AIs themselves become less aligned. Luckily, pretraining LLMs with synthetic data about good AIs helps them become more aligne
saved by
related reading
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentalignmentpretraining.ai
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- [2601.10160] Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentarxiv.org
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Alignment pretraining could backfire — LessWronglesswrong.com
- LLMs Learn to Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactionsarxiv.org
- Model Spec Midtraining: Improving How Alignment Training Generalizesalignment.anthropic.com
- Claude is Now Alignment-Pretrained — LessWronglesswrong.com
- Alignment Faking Mitigationsalignment.anthropic.com
- How far does alignment midtraining generalize?alignment.openai.com
- Teaching Claude Whyalignment.anthropic.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com