The Waluigi Effect (mega-post) - LessWrong
lesswrong.com · 21,929 words · saved by 16 readers
Everyone carries a shadow, and the less it is embodied in the individual’s conscious life, the blacker and denser it is. — Carl Jung …
x The Waluigi Effect (mega-post) — LessWrong Waluigi Effect Simulator Theory ChatGPT Deceptive Alignment Language Models (LLMs) Prompt Engineering RLHF Risks of Astronomical Suffering (S-risks) AI Frontpage 648 The Waluigi Effect (mega-post) by Cleo Nardo 3rd Mar 2023 AI Alignment Forum 19 min read 188 648 Ω 56 Everyone carries a shadow, and the less it is embodied in the individual’s conscious life, the blacker and denser it is. — Carl Jung Acknowlegements: Thanks to Janus and Jozdien for comments. Background In this article, I will present a mechanistic explanation of the Waluigi Effect and
saved by
- Amir
- Svitlana Midianko
- Benjamin Laufer
- brunella k
- Varun Shenoy
- Rona Wang
- Malik Piara
- Adithya V
- Chinhai Hour
- Joey Yap
- Tiffany Trinh
- anka hu
related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- Simulated Users & Sad LLMs1a3orn.com
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Prompt Injection as Role Confusionrole-confusion.github.io
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- The persona selection model — LessWronglesswrong.com
- Towards a Typology of Strange LLM Chains-of-Thought1a3orn.com
- the void — LessWronglesswrong.com