Virtue-Ethical Rationality and Training Dynamics | Manifund
In 'On Eudaimonia and Optimization' I argue that the concept of eudaimonia -- active, rational flourishing -- points to a framework for rational action that differs from standard instrumental planning. In a word, I relate eudaimonic action to commitment to 'promoting x x-ingly' (promoting peace peacefully, promoting mathematics mathematically, promoting democracy democratically), and argue that x-s that can materially sustain such a commitment are an important type of natural abstraction. When turning to AI alignment proper, I argue that: Human values and human ideas of prosocial behavior (what Joe Carlsmith calls 'being nicer than Clippy') become easier to operationalize when considered as eudaimonic practices The 'eudaimonic practices' framework is attuned to RL and RL-related dynamics at multiple scales, from the post-training of frontier models to human sociotechnical dynamics The prospects for instilling AIs with narrowly safety-focussed values such as 'transparency' and possibly
Virtue-Ethical Rationality and Training Dynamics | Manifund 1 Virtue-Ethical Rationality and Training Dynamics Peli Grietzer Active Grant $20,000 raised $40,000 funding goal Donate Sign in to donate In ' On Eudaimonia and Optimization ' I argue that the concept of eudaimonia -- active, rational flourishing -- points to a framework for rational action that differs from standard instrumental planning. In a word, I relate eudaimonic action to commitment to 'promoting x x -ingly' (promoting peace peacefully, promoting mathematics mathematically, promoting democracy democratically), and argue that
Explore this link on the map →related reading
- LessWronglesswrong.com
- LLM Alignment, ethical and mathematical realism, and the most important actions in davidad's understanding — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- [2605.10310] Positive Alignment: Artificial Intelligence for Human Flourishingarxiv.org
- Teaching Claude why \ Anthropicanthropic.com
- The Best of LessWrong — LessWronglesswrong.com
- Aligning to Virtues — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Thomas Larsen's Shortform — LessWronglesswrong.com
- 1a3orn's Shortform — LessWronglesswrong.com
- Synthetic Persona Pretraining: Alignment from Token Zero — LessWronglesswrong.com
- “Alignment Faking” frame is somewhat fake — LessWronglesswrong.com