From personas to intentions: towards a science of motivations for AI models — LessWrong
TLDR: • * Behavior-only descriptions are useful, but insufficient for aligning advanced models with high assurance. * Two models can look equally a…
x From personas to intentions: towards a science of motivations for AI models — LessWrong AI Frontpage 77 From personas to intentions: towards a science of motivations for AI models by David Africa , Jacob Pfau 14th Apr 2026 8 min read 5 77 TLDR: Behavior-only descriptions are useful, but insufficient for aligning advanced models with high assurance. Two models can look equally aligned on ordinary prompts while being driven by very different underlying motivations; this difference may only show up in rare but crucial situations. So persona research should aim to infer motivational structure: t
Explore this link on the map →saved by
related reading
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Emotion concepts and their function in a large language model \ Anthropicanthropic.com
- Clarifying the role of the behavioral selection model — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Your Model Organisms Might Be Fried — LessWronglesswrong.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- The Case for Model Forensics — LessWronglesswrong.com
- Auditing language models for hidden objectives — LessWronglesswrong.com