A gap in the theoretical justification for surrogate goals and safe Pareto improvements – The Universe from an Intentional Stance
The SPI framework tells us that if we choose only between, for instance, aligned delegation and aligned delegation plus surrogate goal, then implementing the surrogate goal is better. This argument in the paper is persuasive pretty much regardless of what kind of beliefs we hold and what notion of rationality we adopt. In particular, it should convince us if we’re expected utility maximizers. However, in general we have more than just these two options (i.e., more than just aligned delegation and aligned delegation plus surrogate goals); we can instruct our delegates in all sorts of ways. The SPI formalism does not directly provide an argument that among all these possible instructions we should implement some instructions that involve surrogate goals. I will call this the surrogate goal justification gap. Can this gap be bridged? If so, what are the necessary and sufficient conditions for bridging the gap? The problem is related to but distinct from other issues with SPIs (such the SP
Short summary and overview The SPI framework tells us that if we choose only between, for instance, aligned delegation and aligned delegation plus surrogate goal, then implementing the surrogate goal is better. This argument in the paper is persuasive pretty much regardless of what kind of beliefs we hold and what notion of rationality we adopt. In particular, it should convince us if we’re expected utility maximizers. However, in general we have more than just these two options (i.e., more than just aligned delegation and aligned delegation plus surrogate goals); we can instruct our delegates
Explore this link on the map →saved by
related reading
- Beliefs are Chosen to Serve Goals — LessWronglesswrong.com
- On Functional Decision Theory :: Wolfgang Schwarzumsu.de
- On The Independence Axiom — LessWronglesswrong.com
- Goal Factoring — LessWronglesswrong.com
- Towards a scale-free theory of intelligent agency — AI Alignment Forumalignmentforum.org
- Optimality is the tiger, and agents are its teeth — LessWronglesswrong.com
- [2512.15584] A Decision-Theoretic Approach for Managing Misalignmentarxiv.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Pascal's mugging - Wikipediaen.wikipedia.org
- [1902.09469] Embedded Agencyarxiv.org
- The Stamp Collector — LessWronglesswrong.com
- UDT shows that decision theory is more puzzling than ever — LessWronglesswrong.com