The behavioral selection model for predicting AI motivations — LessWrong
Highly capable AI systems might end up deciding the future. Understanding what will drive those decisions is therefore one of the most important ques…
x The behavioral selection model for predicting AI motivations — LessWrong Redwood Research Goal-Directedness Goals AI Curated 2025 Top Fifty: 23 % 204 The behavioral selection model for predicting AI motivations by Alex Mallen , Buck 4th Dec 2025 AI Alignment Forum 19 min read 31 204 Ω 82 Highly capable AI systems might end up deciding the future. Understanding what will drive those decisions is therefore one of the most important questions we can ask. Many people have proposed different answers. Some predict that powerful AIs will learn to intrinsically pursue reward. Others respond by sayin
Explore this link on the map →saved by
- Yixiong Hao
- Asher P
- Emil Ryd
- Julian H
- Jo J.
- Arjun Khandelwal
- Kaustubh Kislay
- Arya Pasumarthi
- Vincent Cimino
- 6824
related reading
- From personas to intentions: towards a science of motivations for AI models — LessWronglesswrong.com
- Clarifying the role of the behavioral selection model — LessWronglesswrong.com
- Fitness-Seekers: Generalizing the Reward-Seeking Threat Model — LessWronglesswrong.com
- Reward is not the optimization target — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Beliefs are Chosen to Serve Goals — LessWronglesswrong.com
- Models Don't "Get Reward" — LessWronglesswrong.com
- Aren’t developers regularly making their AIs nice and safe and obedient? | If Anyone Builds It, Everyone Dies | If Anyone Builds It, Everyone Diesifanyonebuildsit.com
- What failure looks like — LessWronglesswrong.com
- Why AIs aren't power-seeking yet — LessWronglesswrong.com
- Why AI alignment could be hard with modern deep learningcold-takes.com