Alignment pretraining could backfire — LessWrong
There has been recent interest in generating synthetic documents to upsample examples of aligned AI during LLM pretraining. See, for instance, Geodes…
x Alignment pretraining could backfire — LessWrong Aligned AI Role-Model Fiction Alignment Pretraining LLM Personas AI Frontpage 43 Alignment pretraining could backfire by Alexandre Variengien 17th Jun 2026 2 min read 9 43 Epistemic status: speculative, but I think the mechanism is plausible. There has been recent interest in generating synthetic documents to upsample examples of aligned AI during LLM pretraining. See, for instance, Geodesic's Alignment Pretraining paper or Anthropic's " Teaching Claude Why ." I worry that this strategy can work well up to moderately capable models but backfir
saved by
related reading
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentalignmentpretraining.ai
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- [2601.10160] Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentarxiv.org
- Claude is Now Alignment-Pretrained — LessWronglesswrong.com
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- Synthetic Persona Pretraining: Alignment from Token Zero — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Synthetic Persona Pretraining: Alignment from Token Zeromodelraising.ai
- How far does alignment midtraining generalize?alignment.openai.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Alignment faking in large language modelsarxiv.org