Claude is Now Alignment-Pretrained — LessWrong
Anthropic are now actively using the approach to alignment often called “Alignment Pretraining” or “Safety Pretraining” — using Stochastic Gradient Descent on a large body of natural or synthetic documents showing the AI assistant doing the right thing in morally challenging situations. They tried this out, found it works well and generalizes well, and they’re now using it. I’m absolutely delighted. I’ve been repeatedly advocating this approach on LessWrong and the Alignment Forum for a couple of years now: I’ve been very excited about this alignment technique ever since I read the seminal paper demonstrating that it was extremely effective: Pretraining Language Models with Human Preferences (Korbak et al., ’23). This was later followed up by Safety Pretraining: Toward the Next Generation of Safe AI (Maini, Goyal, Sam et al., ’25), You Are What You Eat - AI Alignment Requires Understanding How Data Shapes Structure and Generalisation (Lehalleur, Hoogland, Farrugia-Roberts et al., ’25),
x Claude is Now Alignment-Pretrained — LessWrong Aligned AI Role-Model Fiction Alignment Pretraining Anthropic (org) Aligned AI Proposals AI Frontpage 87 Claude is Now Alignment-Pretrained by RogerDearnaley 13th May 2026 2 min read 9 87 This is a linkpost for https://www.anthropic.com/research/teaching-claude-why Anthropic are now actively using the approach to alignment often called “ Alignment Pretraining ” or “Safety Pretraining” — using Stochastic Gradient Descent on a large body of natural or synthetic documents showing the AI assistant doing the right thing in morally challenging situati
Explore this link on the map →saved by
related reading
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- How far does alignment midtraining generalize?alignment.openai.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Alignment pretraining could backfire — LessWronglesswrong.com
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- [2601.10160] Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignmentarxiv.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Synthetic Persona Pretraining: Alignment from Token Zero — LessWronglesswrong.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Alignment faking in large language modelsarxiv.org