Reformative Hypocrisy, and Paying Close Enough Attention to Selectively Reward It. — LessWrong
People often attack frontier AI labs for "hypocrisy" when the labs admit publicly that AI is an extinction threat to humanity. Often these attacks i…
x Reformative Hypocrisy, and Paying Close Enough Attention to Selectively Reward It. — LessWrong AI Frontpage 53 Reformative Hypocrisy, and Paying Close Enough Attention to Selectively Reward It. by Andrew_Critch 11th Sep 2024 3 min read 13 53 People often attack frontier AI labs for "hypocrisy" when the labs admit publicly that AI is an extinction threat to humanity. Often these attacks ignore the difference between various kinds of hypocrisy, some of which are good, including what I'll call "reformative hypocrisy". Attacking good kinds of hypocrisy can be actively harmful for humanity's abil
saved by
related reading
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- The Possessed Machines: Dostoevsky's Demons and the Coming AGI Catastrophepossessedmachines.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Deep Deceptiveness — LessWronglesswrong.com
- A guide to the AI tribes - by Michel Justen - What is thismicheljusten.substack.com
- Why are AI agents lying, cheating and coordinating?yoshuabengio.org
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- What we’d like to fund — Paradigm 3paradigm3.org
- Training a Misaligned Reward Seekeralignment.anthropic.com
- Nicholas Decker In Hellsubstack.com
- The Best of LessWrong — LessWronglesswrong.com