Alignment pretraining could backfire — LessWrong
There has been recent interest in generating synthetic documents to upsample examples of aligned AI during LLM pretraining. See, for instance, Geodes…
x Alignment pretraining could backfire — LessWrong Aligned AI Role-Model Fiction Alignment Pretraining LLM Personas AI Frontpage 43 Alignment pretraining could backfire by Alexandre Variengien 17th Jun 2026 2 min read 9 43 Epistemic status: speculative, but I think the mechanism is plausible. There has been recent interest in generating synthetic documents to upsample examples of aligned AI during LLM pretraining. See, for instance, Geodesic's Alignment Pretraining paper or Anthropic's " Teaching Claude Why ." I worry that this strategy can work well up to moderately capable models but backfir
Explore this link on the map →related reading
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- Claude is Now Alignment-Pretrained — LessWronglesswrong.com
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- Synthetic Persona Pretraining: Alignment from Token Zero — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- [2601.10160] Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentarxiv.org
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- How far does alignment midtraining generalize?alignment.openai.com
- Alignment faking in large language modelsarxiv.org
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- Maybe we should pretrain on synthetic data about good-but-reward-hacking AIs — LessWronglesswrong.com
- AIs Will Increasingly Fake Alignment - by Zvi Mowshowitzthezvi.substack.com