Model Spec Midtraining: Improving How Alignment Training Generalizes
We introduce model spec midtraining (MSM): after pre-training but before alignment fine-tuning, we train models on synthetic documents discussing their Model Spec. This shapes how models generalize from subsequent alignment training. For example, two models with identical alignment fine-tuning can generalize to adopt different values depending on the Model Spec used during MSM. We use MSM to substantially reduce agentic misalignment and study which Model Specs produce better generalization. 📄 Paper, 💻 Code Some frontier AI developers aim to align language models to a Model Spec or Constitution that describes intended model behavior. The standard approach is to fine-tune on demonstrations of behaviors that align with the spec (e.g., conversations where the model acts as intended). However, this can fail to produce robust alignment. For example, LLM agents have been shown to take unethical actions (e.g., blackmailing, leaking company information, alignment faking) when placed in scena
Model Spec Midtraining: Improving How Alignment Training Generalizes Alignment Science Blog Model Spec Midtraining: Improving How Alignment Training Generalizes Chloe Li 1 , Nevan Wichers 1 , May 5, 2026 Sara Price 2 , Samuel Marks 2,† , Jon Kutasov 2,† 1 Anthropic Fellows Program; 2 Anthropic; † Equal advising tl;dr We introduce model spec midtraining (MSM): after pre-training but before alignment fine-tuning, we train models on synthetic documents discussing their Model Spec. This shapes how models generalize from subsequent alignment training. For example, two models with identical alignmen
Explore this link on the map →related reading
- How far does alignment midtraining generalize?alignment.openai.com
- Teaching Claude why \ Anthropicanthropic.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Model Organisms for Emergent Misalignment — LessWronglesswrong.com
- Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment — LessWronglesswrong.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- Alignment faking in large language modelsarxiv.org
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Subliminal Learning: Language Models Transmit Behavioral Traits via Hidden Signals in Dataalignment.anthropic.com
- Thomas Larsen's Shortform — LessWronglesswrong.com