The behavioral selection model for predicting AI motivations
Highly capable AI systems might end up deciding the future. Understanding what will drive those decisions is therefore one of the most important questions we can ask. Many people have proposed different answers. Some predict that powerful AIs will learn to intrinsically pursue reward. Others respond by saying reward is not the optimization target, and instead reward “chisels” a combination of context-dependent cognitive patterns into the AI. Some argue that powerful AIs might end up with an almost arbitrary long-term goal. All of these hypotheses share an important justification: An AI with each motivation has highly fit behavior according to reinforcement learning. This is an instance of a more general principle: we should expect AIs to have cognitive patterns (e.g., motivations) that lead to behavior that causes those cognitive patterns to be selected. In this post I’ll spell out what this more general principle means and why it’s helpful. Specifically: I’ll introduce the “behavioral
Highly capable AI systems might end up deciding the future. Understanding what will drive those decisions is therefore one of the most important questions we can ask. Many people have proposed different answers. Some predict that powerful AIs will learn to intrinsically pursue reward. Others respond by saying reward is not the optimization target, and instead reward “chisels” a combination of context-dependent cognitive patterns into the AI. Some argue that powerful AIs might end up with an almost arbitrary long-term goal. All of these hypotheses share an important justification: An AI…
saved by
related reading
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Fitness-Seekers: Generalizing the Reward-Seeking Threat Model — LessWronglesswrong.com
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- Clarifying the role of the behavioral selection model — LessWronglesswrong.com
- Reward is not the optimization target — LessWronglesswrong.com
- Deep Deceptiveness — LessWronglesswrong.com
- From personas to intentions: towards a science of motivations for AI models — LessWronglesswrong.com
- Models Don't "Get Reward" — LessWronglesswrong.com
- Fail safe(r) at alignment by channeling reward-hacking into a "spillway" motivationblog.redwoodresearch.org
- Reward Is Not the Optimization Targetturntrout.com
- Why AIs aren't power-seeking yet — LessWronglesswrong.com
- Fitness-Seekers: Generalizing the Reward-Seeking Threat Modelsubstack.com