There are no coherence theorems — AI Alignment Forum
[Written by EJT as part of the CAIS Philosophy Fellowship. Thanks to Dan for help posting to the Alignment Forum] For about fifteen years, the AI safety community has been discussing coherence arguments. In papers and posts on the subject, it’s often written that there exist 'coherence theorems' which state that, unless an agent can be represented as maximizing expected utility, that agent is liable to pursue strategies that are dominated by some other available strategy. Despite the prominence of these arguments, authors are often a little hazy about exactly which theorems qualify as coherence theorems. This is no accident. If the authors had tried to be precise, they would have discovered that there are no such theorems. I’m concerned about this. Coherence arguments seem to be a moderately important part of the basic case for existential risk from AI. To spot the error in these arguments, we only have to look up what cited ‘coherence theorems’ actually say. And yet the error seems to
x There are no coherence theorems — AI Alignment Forum CAIS Philosophy Fellowship Midpoint Deliverables Coherence Arguments AI Frontpage 39 There are no coherence theorems by Dan H , Elliott Thornley 20th Feb 2023 23 min read 136 39 [Written by EJT as part of the CAIS Philosophy Fellowship. Thanks to Dan for help posting to the Alignment Forum] Introduction For about fifteen years, the AI safety community has been discussing coherence arguments . In papers and posts on the subject, it’s often written that there exist 'coherence theorems' which state that, unless an agent can be represented as
Explore this link on the map →related reading
- The hot mess theory of AI misalignment: More intelligent agents behave less coherently | Jascha’s blogsohl-dickstein.github.io
- Why The Focus on Expected Utility Maximisers? — LessWronglesswrong.com
- LessWronglesswrong.com
- Towards a scale-free theory of intelligent agency — AI Alignment Forumalignmentforum.org
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- On The Independence Axiom — LessWronglesswrong.com
- Anthropic’s “Hot Mess” paper overstates its case (and the blog post is worse) — LessWronglesswrong.com
- Von Neumann–Morgenstern utility theorem - Wikipediaen.wikipedia.org
- Decision Theory (Stanford Encyclopedia of Philosophy)plato.stanford.edu
- The Best of LessWrong — LessWronglesswrong.com
- Decision Theory FAQ — LessWronglesswrong.com
- Optimality is the tiger, and agents are its teeth — LessWronglesswrong.com