✳flâneur — a map of the web's best reading
The Waluigi Effect (mega-post) - LessWrong
lesswrong.com · 21,929 words · saved by 16 readers
Everyone carries a shadow, and the less it is embodied in the individual’s conscious life, the blacker and denser it is. — Carl Jung …
x The Waluigi Effect (mega-post) — LessWrong Waluigi Effect Simulator Theory ChatGPT Deceptive Alignment Language Models (LLMs) Prompt Engineering RLHF Risks of Astronomical Suffering (S-risks) AI Frontpage 648 The Waluigi Effect (mega-post) by Cleo Nardo 3rd Mar 2023 AI Alignment Forum 19 min read 188 648 Ω 56 Everyone carries a shadow, and the less it is embodied in the individual’s conscious life, the blacker and denser it is. — Carl Jung Acknowlegements: Thanks to Janus and Jozdien for comments. Background In this article, I will present a mechanistic explanation of the Waluigi Effect and
Explore this link on the map →saved by
- Amir
- Svitlana Midianko
- Benjamin Laufer
- brunella k
- Varun Shenoy
- Rona Wang
- Malik Piara
- Adithya V
- Chinhai Hour
- Joey Yap
- Tiffany Trinh
- anka hu
related reading
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- RLHF: Reinforcement Learning from Human Feedbackhuyenchip.com
- Disagreeable Me: A Multi-Level view of LLM Intentionalitydisagreeableme.blogspot.com
- the void — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- The persona selection model — LessWronglesswrong.com
- Alignment will happen by default. What’s next? — LessWronglesswrong.com
- Taking LLMs Seriously (As Language Models) — LessWronglesswrong.com