The Artificiality of Alignment - by jessica dai - Reboot
joinreboot.org · 4,105 words · saved by 7 readers
How are we actually “aligning AI with human values”?
Long essay today; let’s get straight to it. (Note — you might have a better time with footnotes on web/app than email, if that matters to you.) Nature by Alan Warburton, CC 4.0 By Jessica Dai Credulous, breathless coverage of “AI existential risk” (abbreviated “x-risk”) has reached the mainstream. Who could have foreseen that the smallcaps onomatopoeia “ꜰᴏᴏᴍ” — both evocative of and directly derived from children’s cartoons — might show up uncritically in the New Yorker? More than ever, the public discourse about AI and its risks, and about what can or should be done about those risks, is…
saved by
related reading
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Value Alignment Is a Pseudo Conceptzilanqian.substack.com
- Existential Risk from AI: An Exposition for Mathematiciansalkjash.github.io
- Core views on AI safety: When, why, what, and how \ Anthropicanthropic.com
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- Leaving Open Philanthropy, going to Anthropic - Joe Carlsmithjoecarlsmith.com
- Where I agree and disagree with Eliezer — LessWronglesswrong.com
- What is AI alignment? - by Adam Jones - BlueDot Impactblog.bluedot.org
- The case against AI alignment — LessWronglesswrong.com
- [2605.10310] Positive Alignment: Artificial Intelligence for Human Flourishingarxiv.org