Deep Deceptiveness — LessWrong
There are some obvious ways you might try to train deceptiveness out of AIs. But deceptiveness can emerge from the recombination of non-deceptive cog…
x Deep Deceptiveness — LessWrong Best of LessWrong 2023 Deception Deceptive Alignment Threat Models (AI) AI Frontpage 283 Deep Deceptiveness by So8res 21st Mar 2023 AI Alignment Forum 17 min read 60 283 Ω 95 Meta This post is an attempt to gesture at a class of AI notkilleveryoneism (alignment) problem that seems to me to go largely unrecognized. E.g., it isn’t discussed (or at least I don't recognize it) in the recent plans written up by OpenAI ( 1 , 2 ), by DeepMind’s alignment team , or by Anthropic , and I know of no other acknowledgment of this issue by major labs. You could think of this
Explore this link on the map →related reading
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- How confessions can keep language models honest | OpenAIopenai.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Can Agents Fool Each Other? - AI Villagetheaidigest.org
- Interpretability Will Not Reliably Find Deceptive AI — LessWronglesswrong.com
- Ajeya Cotra on accidentally teaching AI models to deceive us | 80,000 Hours80000hours.org
- Natural Deception with RL - Rajan Agarwalrajan.sh
- The Road To Honest AI - by Scott Alexanderastralcodexten.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com