Alignment & Succession: The Two Bars of Alignment
nosetgauge.com · 4,185 words · saved by 2 readers
Why much AI alignment discourse automatically treats the transition to powerful AI as a succession problem
Rembrandt, Saul and David When people talk about AI being “aligned”, I think there’s a lot of conflation between two different bars of success: The AI does what you say and does not go rogue. If you ask it to do a machine learning experiment for you, it actually does that instead of scheming against you and escaping onto the internet. If you say “stop”, it stops. (Some AI properties considered important for achieving this bar are corrigibility, non-deceptiveness, and intent alignment.) The AI has fully internalized our values, and could run society, and the lives of the humans in it, in a…
saved by
related reading
- Alignment & Succession: The Ideology of Successionnosetgauge.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- What failure looks like — LessWronglesswrong.com
- The Artificiality of Alignmentjoinreboot.org
- What failure looks like — AI Alignment Forumalignmentforum.org
- The Problem — LessWronglesswrong.com
- The case against AI alignment — LessWronglesswrong.com
- Value Alignment Is a Pseudo Conceptzilanqian.substack.com
- What is AI alignment? - by Adam Jones - BlueDot Impactblog.bluedot.org
- Alignment & Succession: Morality Lives in the Human Individualnosetgauge.com
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org