Aligning to Virtues — LessWrong
Which alignment target? Suppose you’re an AI company or government, and you want to figure out what values to align your AI to. Here are three option…
x Aligning to Virtues — LessWrong AI Frontpage 93 Aligning to Virtues by Richard_Ngo 16th Feb 2026 5 min read 36 93 Which alignment target? Suppose you’re an AI company or government, and you want to figure out what values to align your AI to. Here are three options, and some of their downsides: AIs that are aligned to a set of consequentialist values are incentivized to acquire power to pursue those values. This creates power struggles between those AIs and: Humans who don’t share those values. Humans who disagree with the AI about how to pursue those values. Humans who don’t trust that the A
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- [2605.10310] Positive Alignment: Artificial Intelligence for Human Flourishingarxiv.org
- The case against AI alignment — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Deontology and virtue ethics as "effective theories" of consequentialist ethics — LessWronglesswrong.com
- Virtue Ethics for Consequentialists — LessWronglesswrong.com
- What Is The Alignment Problem? — LessWronglesswrong.com
- “Alignment Faking” frame is somewhat fake — LessWronglesswrong.com
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- LLM Alignment, ethical and mathematical realism, and the most important actions in davidad's understanding — LessWronglesswrong.com
- The Machines Lack Honour — LessWronglesswrong.com