Fitness-Seekers: Generalizing the Reward-Seeking Threat Model — LessWrong
lesswrong.com · 8,104 words · saved by 1 readers
If you think reward-seekers are plausible, you should also think “fitness-seekers” are plausible. But their risks aren't the same. …
x Fitness-Seekers: Generalizing the Reward-Seeking Threat Model — LessWrong Fitness-seeking AIs Outer Alignment Redwood Research AI Frontpage 92 Fitness-Seekers: Generalizing the Reward-Seeking Threat Model by Alex Mallen 29th Jan 2026 AI Alignment Forum 21 min read 5 92 Ω 44 If you think reward-seekers are plausible, you should also think “fitness-seekers” are plausible. But their risks aren't the same. The AI safety community often emphasizes reward-seeking as a central case of a misaligned AI alongside scheming (e.g., Cotra’s sycophant vs schemer, Carlsmith’s terminal vs instrumental traini
saved by
related reading
- The behavioral selection model for predicting AI motivationsblog.redwoodresearch.org
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Fitness-Seekers: Generalizing the Reward-Seeking Threat Modelsubstack.com
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- Risk from fitness-seeking AIs: mechanisms and mitigationssubstack.com
- Reward is not the optimization target — LessWronglesswrong.com
- Risk from fitness-seeking AIs: mechanisms and mitigationsblog.redwoodresearch.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Are AIs more likely to pursue on-episode or beyond-episode reward?blog.redwoodresearch.org
- What failure looks like — LessWronglesswrong.com
- Measuring Reward-Seeking by Instilling Contrastive Beliefsalignment.openai.com
- What failure looks like — AI Alignment Forumalignmentforum.org