✳flâneur — a map of the web's best reading
A Three-Layer Model of LLM Psychology — LessWrong
lesswrong.com · 6,393 words · saved by 1 readers
This post offers an accessible model of psychology of character-trained LLMs like Claude. …
x A Three-Layer Model of LLM Psychology — LessWrong Best of LessWrong 2024 LLM Personas AI Psychology Aligned AI Role-Model Fiction Alignment Pretraining Cyborgism Language model cognitive architecture AI Curated 266 A Three-Layer Model of LLM Psychology by Jan_Kulveit 26th Dec 2024 AI Alignment Forum 10 min read 17 266 Ω 84 This post offers an accessible model of psychology of character-trained LLMs like Claude. Epistemic Status This is primarily a phenomenological model based on extensive interactions with LLMs, particularly Claude. It's intentionally anthropomorphic in cases where I believe
Explore this link on the map →related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- The persona selection model — LessWronglesswrong.com
- Role-playing vs Self-modelling — LessWronglesswrong.com
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- AI Induced Psychosis: A shallow investigation — LessWronglesswrong.com
- Claude’s Character \ Anthropicanthropic.com
- Tracing the thoughts of a large language model \ Anthropicanthropic.com
- the void — LessWronglesswrong.com
- On the functional self of LLMs — LessWronglesswrong.com