The persona selection model — LessWrong
TL;DR We describe the persona selection model (PSM): the idea that LLMs learn to simulate diverse characters during pre-training, and post-training e…
x The persona selection model — LessWrong AI Frontpage 2026 Top Fifty: 14 % 176 The persona selection model by Sam Marks 23rd Feb 2026 AI Alignment Forum Linkpost for alignment.anthropic.com 52 min read 53 176 Ω 74 TL;DR We describe the persona selection model (PSM): the idea that LLMs learn to simulate diverse characters during pre-training, and post-training elicits and refines a particular such Assistant persona. Interactions with an AI assistant are then well-understood as being interactions with the Assistant—something roughly like a character in an LLM-generated story. We survey empirica
Explore this link on the map →saved by
related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- The persona selection model \ Anthropicanthropic.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- Role-playing vs Self-modelling — LessWronglesswrong.com
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- the void — LessWronglesswrong.com
- A Case for Model Persona Research — LessWronglesswrong.com
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- [2601.10387] The Assistant Axis: Situating and Stabilizing the Default Persona of Language Modelsarxiv.org
- The assistant axis \ Anthropicanthropic.com