The behavioral selection model for predicting AI motivations — LessWrong
Highly capable AI systems might end up deciding the future. Understanding what will drive those decisions is therefore one of the most important ques…
x The behavioral selection model for predicting AI motivations — LessWrong Redwood Research Goal-Directedness Goals AI Curated 2025 Top Fifty: 23 % 204 The behavioral selection model for predicting AI motivations by Alex Mallen , Buck 4th Dec 2025 AI Alignment Forum 19 min read 31 204 Ω 82 Highly capable AI systems might end up deciding the future. Understanding what will drive those decisions is therefore one of the most important questions we can ask. Many people have proposed different answers. Some predict that powerful AIs will learn to intrinsically pursue reward. Others respond by sayin
saved by
- Yixiong Hao
- Asher P
- Lydia Nottingham
- Emil Ryd
- Julian H
- Jo J.
- Arjun Khandelwal
- Kaustubh Kislay
- Will Anderson
- Arya Pasumarthi
- Vincent Cimino
- Skye
related reading
- The behavioral selection model for predicting AI motivationsblog.redwoodresearch.org
- Clarifying the role of the behavioral selection model — LessWronglesswrong.com
- From personas to intentions: towards a science of motivations for AI models — LessWronglesswrong.com
- Reward is not the optimization target — LessWronglesswrong.com
- Fitness-Seekers: Generalizing the Reward-Seeking Threat Model — LessWronglesswrong.com
- Fail safe(r) at alignment by channeling reward-hacking into a "spillway" motivationblog.redwoodresearch.org
- Deep Deceptiveness — LessWronglesswrong.com
- Measuring Reward-Seeking by Instilling Contrastive Beliefsalignment.openai.com
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- What failure looks like — AI Alignment Forumalignmentforum.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Risk from fitness-seeking AIs: mechanisms and mitigationssubstack.com