A Case for Model Persona Research — LessWrong
Context: At the Center on Long-Term Risk (CLR) our empirical research agenda focuses on studying (malicious) personas, their relation to generalizati…
x A Case for Model Persona Research — LessWrong LLM Personas AI Frontpage 2025 Top Fifty: 15 % 121 A Case for Model Persona Research by nielsrolf , Maxime Riché , Daniel Tan 15th Dec 2025 5 min read 11 121 Context: At the Center on Long-Term Risk (CLR) our empirical research agenda focuses on studying (malicious) personas, their relation to generalization, and how to prevent misgeneralization, especially given weak overseers (e.g., undetected reward hacking) or underspecified training signals. This has motivated our past research on Emergent Misalignment and Inoculation Prompting , and we want
Explore this link on the map →related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- The persona selection model — LessWronglesswrong.com
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- Chapter 4: Alignment Science - ARENAlearn.arena.education
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- the void — LessWronglesswrong.com
- Persona vectors: Monitoring and controlling character traits in language models \ Anthropicanthropic.com
- The persona selection model \ Anthropicanthropic.com
- Synthetic Persona Pretraining: Alignment from Token Zero — LessWronglesswrong.com
- Role-playing vs Self-modelling — LessWronglesswrong.com