The Persona Selection Model: Why AI Assistants might Behave like Humans
We describe the persona selection model (PSM): the idea that LLMs learn to simulate diverse characters during pre-training, and post-training elicits and refines a particular such Assistant persona. Interactions with an AI assistant are then well-understood as being interactions with the Assistant—something roughly like a character in an LLM-generated story. We survey empirical behavioral, generalization, and interpretability-based evidence for PSM. PSM has consequences for AI development, such as recommending anthropomorphic reasoning about AI psychology and introduction of positive AI archetypes into training data. An important open question is how exhaustive PSM is, especially whether there might be sources of agency external to the Assistant persona, and how this might change in the future. What sort of thing is a modern AI assistant? One perspective holds that they are shallow, rigid systems that narrowly pattern-match user inputs to training data. Another perspective regards AI s
The Persona Selection Model: Why AI Assistants might Behave like Humans Alignment Science Blog The Persona Selection Model: Why AI Assistants might Behave like Humans Sam Marks, Jack Lindsey, Christopher Olah February 23, 2026 tl;dr We describe the persona selection model (PSM): the idea that LLMs learn to simulate diverse characters during pre-training, and post-training elicits and refines a particular such Assistant persona. Interactions with an AI assistant are then well-understood as being interactions with the Assistant—something roughly like a character in an LLM-generated story. We sur
Explore this link on the map →saved by
- Freeman Jiang
- Grace Gong
- Asher P
- Aaron Pham
- Laerdon Kim
- Lydia Nottingham
- Sudarsh K
- Yudhister Joel Kumar
- Timothy Kostolansky
- Sarah
- Dhruv Sheth
- Eric Huang
related reading
- The persona selection model — LessWronglesswrong.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Role-playing vs Self-modelling — LessWronglesswrong.com
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- The persona selection model \ Anthropicanthropic.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Claude’s Character \ Anthropicanthropic.com
- AI Induced Psychosis: A shallow investigation — LessWronglesswrong.com
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- The assistant axis \ Anthropicanthropic.com
- [2601.10387] The Assistant Axis: Situating and Stabilizing the Default Persona of Language Modelsarxiv.org