After Orthogonality: Virtue-Ethical Agency and AI Alignment
Preface This essay argues that rational people don’t have goals, and that rational AIs shouldn’t have goals. Human actions are rational not because we direct them at some final ‘goals,’ but because we align actions to practices[1]: networks of actions, action-dispositions, action-evaluation criteria, and action-resources that structure,
Preface This essay argues that rational people don’t have goals, and that rational AIs shouldn’t have goals. Human actions are rational not because we direct them at some final ‘goals,’ but because we align actions to practices[1]: networks of actions, action-dispositions, action-evaluation criteria, and action-resources that structure, clarify, develop, and promote themselves. If we want AIs that can genuinely support, collaborate with, or even comply with human agency, AI agents’ deliberations must share a “type signature” with the practices-based logic we use to reflect and act. I argue…
saved by
related reading
- Value Alignment Is a Pseudo Conceptzilanqian.substack.com
- LessWronglesswrong.com
- Virtue-Ethical Rationality and Training Dynamics | Manifundmanifund.org
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- [2605.10310] Positive Alignment: Artificial Intelligence for Human Flourishingarxiv.org
- [2001.09768] Artificial Intelligence, Values and Alignmentarxiv.org
- The Best of LessWrong — LessWronglesswrong.com
- Beliefs are Chosen to Serve Goals — LessWronglesswrong.com
- The Artificiality of Alignmentjoinreboot.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- AI Goals Forecast — AI 2027ai-2027.com
- Aligning to Virtues — LessWronglesswrong.com