Stress-testing model specs reveals character differences among language models
We generate over 300,000 user queries that trade-off value-based principles in model specifications. Under these scenarios, we observe distinct value prioritization and behavior patterns in frontier models from Anthropic, OpenAI, Google DeepMind and xAI. Our experiments also uncovered thousands of cases of direct contradictions or interpretive ambiguities within the model spec. 📄 Paper, 📊 Dataset Model specifications are the behavioral guidelines that large language models are trained to follow. They list principles like "be helpful," "assume good intentions," or "stay within safety bounds." Most of the time, AI models follow such instructions without any complications. But what happens when these principles clash? Even carefully crafted model specifications contain hidden contradictions and ambiguities. In a new paper, led by participants in the Anthropic Fellows program and in collaboration with researchers at the Thinking Machines Lab, we expose these ‘specification gaps’ by gener
Stress-testing model specs reveals character differences among language models Alignment Science Blog Stress-testing model specs reveals character differences among language models Jifan Zhang 1 , Henry Sleight 2 , Andi Peng 3 , John Schulman 4 , Esin Durmus 3 October 24, 2025 1 Anthropic Fellows Program; 2 Constellation; 3 Anthropic; 4 Thinking Machines tl;dr We generate over 300,000 user queries that trade-off value-based principles in model specifications. Under these scenarios, we observe distinct value prioritization and behavior patterns in frontier models from Anthropic, OpenAI, Google
related reading
- [2510.07686] Stress-Testing Model Specs Reveals Character Differences among Language Modelsarxiv.org
- Teaching Claude Whyalignment.anthropic.com
- [2605.24229] How Well Do Models Follow Their Constitutions?arxiv.org
- How Claude's values vary by model and language \ Anthropicanthropic.com
- How well do models follow their constitutions? — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Claude Opus 4.5: Model Card, Alignment and Safetythezvi.substack.com
- (a) Anthropic constitution.arxiv.org
- Model Spec Midtraining: Improving How Alignment Training Generalizesalignment.anthropic.com
- [2607.14345] Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Valuesarxiv.org