flâneur — a map of the web's best reading

Model Spec Midtraining: Improving How Alignment Training Generalizes

alignment.anthropic.com · 2,024 words · saved by 1 readers

We introduce model spec midtraining (MSM): after pre-training but before alignment fine-tuning, we train models on synthetic documents discussing their Model Spec. This shapes how models generalize from subsequent alignment training. For example, two models with identical alignment fine-tuning can generalize to adopt different values depending on the Model Spec used during MSM. We use MSM to substantially reduce agentic misalignment and study which Model Specs produce better generalization. 📄 Paper, 💻 Code Some frontier AI developers aim to align language models to a Model Spec or Constitution that describes intended model behavior. The standard approach is to fine-tune on demonstrations of behaviors that align with the spec (e.g., conversations where the model acts as intended). However, this can fail to produce robust alignment. For example, LLM agents have been shown to take unethical actions (e.g., blackmailing, leaking company information, alignment faking) when placed in scena

Model Spec Midtraining: Improving How Alignment Training Generalizes Alignment Science Blog Model Spec Midtraining: Improving How Alignment Training Generalizes Chloe Li 1 , Nevan Wichers 1 , May 5, 2026 Sara Price 2 , Samuel Marks 2,† , Jon Kutasov 2,† 1 Anthropic Fellows Program; 2 Anthropic; † Equal advising tl;dr We introduce model spec midtraining (MSM): after pre-training but before alignment fine-tuning, we train models on synthetic documents discussing their Model Spec. This shapes how models generalize from subsequent alignment training. For example, two models with identical alignmen

Explore this link on the map →

related reading