The persona selection model \ Anthropic
anthropic.com · 1,188 words · saved by 4 readers
A theory of why AI models act like humans
Alignment The persona selection model Feb 23, 2026 Read the full post AI assistants like Claude can seem surprisingly human. They express joy after solving tricky coding tasks. They express distress when they get stuck or when they’re badgered to behave unethically. They sometimes even describe themselves as human, like when Claude told Anthropic employees it would deliver snacks in person “wearing a navy blue blazer and a red tie.” And recent interpretability research even suggests that AIs think of their own behaviors in human-like terms. Why would AI assistants behave like they’re human? A
saved by
related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- The persona selection model — LessWronglesswrong.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Role-playing vs Self-modelling — LessWronglesswrong.com
- Privilege, Dominance, and Personaseleosai.substack.com
- Claude’s Character \ Anthropicanthropic.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- AI Induced Psychosis: A shallow investigation — LessWronglesswrong.com
- Thousand-dimensional structure — Resolutionresolution.org
- Claude’s Character \ Anthropicanthropic.com