flâneur — a map of the web's best reading

Deep Deceptiveness — LessWrong

lesswrong.com · 22,149 words · saved by 1 readers

There are some obvious ways you might try to train deceptiveness out of AIs. But deceptiveness can emerge from the recombination of non-deceptive cog…

x Deep Deceptiveness — LessWrong Best of LessWrong 2023 Deception Deceptive Alignment Threat Models (AI) AI Frontpage 283 Deep Deceptiveness by So8res 21st Mar 2023 AI Alignment Forum 17 min read 60 283 Ω 95 Meta This post is an attempt to gesture at a class of AI notkilleveryoneism (alignment) problem that seems to me to go largely unrecognized. E.g., it isn’t discussed (or at least I don't recognize it) in the recent plans written up by OpenAI ( 1 , 2 ), by DeepMind’s alignment team , or by Anthropic , and I know of no other acknowledgment of this issue by major labs. You could think of this

Explore this link on the map →

related reading