Deep Deceptiveness — LessWrong
There are some obvious ways you might try to train deceptiveness out of AIs. But deceptiveness can emerge from the recombination of non-deceptive cog…
x Deep Deceptiveness — LessWrong Best of LessWrong 2023 Deception Deceptive Alignment Threat Models (AI) AI Frontpage 283 Deep Deceptiveness by So8res 21st Mar 2023 AI Alignment Forum 17 min read 60 283 Ω 95 Meta This post is an attempt to gesture at a class of AI notkilleveryoneism (alignment) problem that seems to me to go largely unrecognized. E.g., it isn’t discussed (or at least I don't recognize it) in the recent plans written up by OpenAI ( 1 , 2 ), by DeepMind’s alignment team , or by Anthropic , and I know of no other acknowledgment of this issue by major labs. You could think of this
saved by
related reading
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- Self-Other Overlap: A Neglected Approach to AI Alignment — LessWronglesswrong.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- How AI Is Learning to Think in Secretnickandresen.substack.com
- Why We Are Excited About Confessionsalignment.openai.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- [2602.15515] The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probesarxiv.org
- Why are AI agents lying, cheating and coordinating?yoshuabengio.org
- On Anthropic's Sleeper Agents Paperthezvi.substack.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org