flâneur — a map of the web's best reading

The Persona Selection Model: Why AI Assistants might Behave like Humans

alignment.anthropic.com · 13,586 words · saved by 28 readers

We describe the persona selection model (PSM): the idea that LLMs learn to simulate diverse characters during pre-training, and post-training elicits and refines a particular such Assistant persona. Interactions with an AI assistant are then well-understood as being interactions with the Assistant—something roughly like a character in an LLM-generated story. We survey empirical behavioral, generalization, and interpretability-based evidence for PSM. PSM has consequences for AI development, such as recommending anthropomorphic reasoning about AI psychology and introduction of positive AI archetypes into training data. An important open question is how exhaustive PSM is, especially whether there might be sources of agency external to the Assistant persona, and how this might change in the future. What sort of thing is a modern AI assistant? One perspective holds that they are shallow, rigid systems that narrowly pattern-match user inputs to training data. Another perspective regards AI s

The Persona Selection Model: Why AI Assistants might Behave like Humans Alignment Science Blog The Persona Selection Model: Why AI Assistants might Behave like Humans Sam Marks, Jack Lindsey, Christopher Olah February 23, 2026 tl;dr We describe the persona selection model (PSM): the idea that LLMs learn to simulate diverse characters during pre-training, and post-training elicits and refines a particular such Assistant persona. Interactions with an AI assistant are then well-understood as being interactions with the Assistant—something roughly like a character in an LLM-generated story. We sur

Explore this link on the map →

saved by

related reading