[2510.07686] Stress-Testing Model Specs Reveals Character Differences among Language Models
YOU make open access possible! Tell us why you support #openaccess and give to arXiv this week to help keep science open for all. Help | Advanced Search arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website. Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them. Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs. arXiv Operational Status
[2510.07686] Stress-Testing Model Specs Reveals Character Differences among Language Models Skip to main content arXiv is now an independent nonprofit! Learn more × Search arXiv Press Enter to search · Advanced search --> Computer Science > Computation and Language arXiv:2510.07686 (cs) [Submitted on 9 Oct 2025 ( v1 ), last revised 23 Oct 2025 (this version, v2)] Title: Stress-Testing Model Specs Reveals Character Differences among Language Models Authors: Jifan Zhang , Henry Sleight , Andi Peng , John Schulman , Esin Durmus View a PDF of the paper titled Stress-Testing Model Spec
Explore this link on the map →related reading
- Stress-testing model specs reveals character differences among language modelsalignment.anthropic.com
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- Emotion Concepts and their Function in a Large Language Modeltransformer-circuits.pub
- 2025: The year in LLMssimonwillison.net
- How well do models follow their constitutions? — LessWronglesswrong.com
- (a) Anthropic constitution.arxiv.org
- A “diff” tool for AI: Finding behavioral differences in new models \ Anthropicanthropic.com
- The bitter lesson of LLM evalsparsed.com
- Surfacing Pathological Behaviors in Language Models | Transluce AItransluce.org