Stress-testing model specs reveals character differences among language models
We generate over 300,000 user queries that trade-off value-based principles in model specifications. Under these scenarios, we observe distinct value prioritization and behavior patterns in frontier models from Anthropic, OpenAI, Google DeepMind and xAI. Our experiments also uncovered thousands of cases of direct contradictions or interpretive ambiguities within the model spec. 📄 Paper, 📊 Dataset Model specifications are the behavioral guidelines that large language models are trained to follow. They list principles like "be helpful," "assume good intentions," or "stay within safety bounds." Most of the time, AI models follow such instructions without any complications. But what happens when these principles clash? Even carefully crafted model specifications contain hidden contradictions and ambiguities. In a new paper, led by participants in the Anthropic Fellows program and in collaboration with researchers at the Thinking Machines Lab, we expose these ‘specification gaps’ by gener
Stress-testing model specs reveals character differences among language models Alignment Science Blog Stress-testing model specs reveals character differences among language models Jifan Zhang 1 , Henry Sleight 2 , Andi Peng 3 , John Schulman 4 , Esin Durmus 3 October 24, 2025 1 Anthropic Fellows Program; 2 Constellation; 3 Anthropic; 4 Thinking Machines tl;dr We generate over 300,000 user queries that trade-off value-based principles in model specifications. Under these scenarios, we observe distinct value prioritization and behavior patterns in frontier models from Anthropic, OpenAI, Google
Explore this link on the map →related reading
- [2510.07686] Stress-Testing Model Specs Reveals Character Differences among Language Modelsarxiv.org
- Teaching Claude why \ Anthropicanthropic.com
- How well do models follow their constitutions? — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- Thoughts on Claude’s Constitution – Windows On Theorywindowsontheory.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Claude 4 System Cardwww-cdn.anthropic.com
- (a) Anthropic constitution.arxiv.org
- How Claude's values vary by model and language \ Anthropicanthropic.com
- Claude’s Character \ Anthropicanthropic.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org