flâneur — a map of the web's best reading

Stress-testing model specs reveals character differences among language models

alignment.anthropic.com · 2,528 words · saved by 1 readers

We generate over 300,000 user queries that trade-off value-based principles in model specifications. Under these scenarios, we observe distinct value prioritization and behavior patterns in frontier models from Anthropic, OpenAI, Google DeepMind and xAI. Our experiments also uncovered thousands of cases of direct contradictions or interpretive ambiguities within the model spec. 📄 Paper, 📊 Dataset Model specifications are the behavioral guidelines that large language models are trained to follow. They list principles like "be helpful," "assume good intentions," or "stay within safety bounds." Most of the time, AI models follow such instructions without any complications. But what happens when these principles clash? Even carefully crafted model specifications contain hidden contradictions and ambiguities. In a new paper, led by participants in the Anthropic Fellows program and in collaboration with researchers at the Thinking Machines Lab, we expose these ‘specification gaps’ by gener

Stress-testing model specs reveals character differences among language models Alignment Science Blog Stress-testing model specs reveals character differences among language models Jifan Zhang 1 , Henry Sleight 2 , Andi Peng 3 , John Schulman 4 , Esin Durmus 3 October 24, 2025 1 Anthropic Fellows Program; 2 Constellation; 3 Anthropic; 4 Thinking Machines tl;dr We generate over 300,000 user queries that trade-off value-based principles in model specifications. Under these scenarios, we observe distinct value prioritization and behavior patterns in frontier models from Anthropic, OpenAI, Google

Explore this link on the map →

related reading