Risk-Averse AIs
We argue that training AIs to be risk-averse – to treat resources as having diminishing marginal utility – could both preserve AIs’ usefulness (if they turn out aligned) and provide an extra line of defense (if they turn out misaligned).
23rd June 2026 [AI Narration] Risk-Averse AIs Playback speed Volume 0:00 of 1:56:03 Abstract We make the case for training AIs to be risk-averse in resources — specifically, to treat resources as having diminishing marginal utility. These AIs would (for example) choose $40 for sure over a half-chance of $100 and a half-chance of $0. We argue that risk aversion can preserve AIs’ usefulness in the event that they turn out aligned, and that it provides an extra line of defense in the event that AIs turn out misaligned: misaligned but risk-averse AIs would prefer a higher chance of modest…
saved by
related reading
- Can risk aversion learned at low stakes generalize to astronomically high stakes?substack.com
- Many arguments for AI x-risk are wrong — AI Alignment Forumalignmentforum.org
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- Why are AI agents lying, cheating and coordinating?yoshuabengio.org
- Existential Risk from AI: An Exposition for Mathematiciansalkjash.github.io
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Dario Amodei — The Adolescence of Technologydarioamodei.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Reading Listblog.redwoodresearch.org
- Plans A, B, C, and D for misalignment risk — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org