the void — LessWrong
A long essay about LLMs, the nature and history of the the HHH assistant persona, and the implications for alignment. Multiple people have asked me whether I could post this LW in some form, hence this linkpost. ~17,000 words. Originally written on June 7, 2025. (Note: although I expect this post will be interesting to people on LW, keep in mind that it was written with a broader audience in mind than my posts and comments here. This had various implications about my choices of presentation and tone, about which things I explained from scratch rather than assuming as background, my level of comfort casually reciting factual details from memory rather than explicitly checking them against the original source, etc. Although, come of think of it, this was also true of most of my early posts on LW [which were crossposts from my blog], so maybe it's not a big deal...) In particular, your argument that putting material into the world about LLMs potentially becoming misaligned may cause prob
x the void — LessWrong Language Models (LLMs) LLM Personas Hyperstitions Aligned AI Role-Model Fiction Self Fulfilling/Refuting Prophecies AI Frontpage 2025 Top Fifty: 72 % 426 the void by nostalgebraist 11th Jun 2025 AI Alignment Forum 1 min read 108 426 Ω 106 This is a linkpost for https://nostalgebraist.tumblr.com/post/785766737747574784/the-void A long essay about LLMs, the nature and history of the the HHH assistant persona, and the implications for alignment. Multiple people have asked me whether I could post this LW in some form, hence this linkpost. ~17,000 words. Originally written on
Explore this link on the map →related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- The persona selection model — LessWronglesswrong.com
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- Alignment will happen by default. What’s next? — LessWronglesswrong.com
- AI in 2025: gestalt — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- The Owned Ones — LessWronglesswrong.com