Dylan HadfieldMenell on X: "i'm dylan hadfield-menell and this is what i most want to say --- has to be about ai alignment obviously so the standard story: you have an ai, you give it an objective, it optimizes the objective, and if the objective is even slightly wrong at scale you get catastrophe." / X
i'm dylan hadfield-menell and this is what i most want to say --- has to be about ai alignment obviously so the standard story: you have an ai, you give it an objective, it optimizes the objective, and if the objective is even slightly wrong at scale you get catastrophe.
i'm dylan hadfield-menell and this is what i most want to say --- has to be about ai alignment obviously so the standard story: you have an ai, you give it an objective, it optimizes the objective, and if the objective is even slightly wrong at scale you get catastrophe. paperclips, king midas, the monkey's paw. and the proposed solution is: get the objective right. specify what we actually want. and i keep coming back to: that framing is the mistake. not the difficulty of the specification problem — the assumption that a specification is the right kind of object to be building toward at…
saved by
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- What is AI alignment? - by Adam Jones - BlueDot Impactblog.bluedot.org
- The Artificiality of Alignmentjoinreboot.org
- What failure looks like — AI Alignment Forumalignmentforum.org
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Shah and Yudkowsky on alignment failures — LessWronglesswrong.com
- What Is The Alignment Problem? — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Obedient AI - Nina Panicksseryblog.ninapanickssery.com