Persona vectors: Monitoring and controlling character traits in language models \ Anthropic
Language models are strange beasts. In many ways they appear to have human-like “personalities” and “moods,” but these traits are highly fluid and liable to change unexpectedly. Sometimes these changes are dramatic. In 2023, Microsoft's Bing chatbot famously adopted an alter-ego called "Sydney,” which declared love for users and made threats of blackmail. More recently, xAI’s Grok chatbot would for a brief period sometimes identify as “MechaHitler” and make antisemitic comments. Other personality changes are subtler but still unsettling, like when models start sucking up to users or making up facts. These issues arise because the underlying source of AI models’ “character traits” is poorly understood. At Anthropic, we try to shape our models’ characteristics in positive ways, but this is more of an art than a science. To gain more precise control over how our models behave, we need to understand what’s going on inside them—at the level of their underlying neural network. In a new paper
Interpretability Persona vectors: Monitoring and controlling character traits in language models Aug 1, 2025 Read the paper Language models are strange beasts. In many ways they appear to have human-like “personalities” and “moods,” but these traits are highly fluid and liable to change unexpectedly. Sometimes these changes are dramatic. In 2023, Microsoft's Bing chatbot famously adopted an alter-ego called "Sydney,” which declared love for users and made threats of blackmail . More recently, xAI’s Grok chatbot would for a brief period sometimes identify as “MechaHitler” and make antisemitic c
Explore this link on the map →related reading
- [2507.21509] Persona Vectors: Monitoring and Controlling Character Traits in Language Modelsarxiv.org
- [2507.21509] Persona Vectors: Monitoring and Controlling Character Traits in Language Modelsarxiv.org
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- [2601.10387] The Assistant Axis: Situating and Stabilizing the Default Persona of Language Modelsarxiv.org
- The persona selection model — LessWronglesswrong.com
- The assistant axis \ Anthropicanthropic.com
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- [2601.10387] The Assistant Axis: Situating and Stabilizing the Default Persona of Language Modelsarxiv.org
- Claude’s Character \ Anthropicanthropic.com