✳flâneur — a map of the web's best reading
Fitness-Seekers: Generalizing the Reward-Seeking Threat Model — LessWrong
lesswrong.com · 8,104 words · saved by 1 readers
If you think reward-seekers are plausible, you should also think “fitness-seekers” are plausible. But their risks aren't the same. …
x Fitness-Seekers: Generalizing the Reward-Seeking Threat Model — LessWrong Fitness-seeking AIs Outer Alignment Redwood Research AI Frontpage 92 Fitness-Seekers: Generalizing the Reward-Seeking Threat Model by Alex Mallen 29th Jan 2026 AI Alignment Forum 21 min read 5 92 Ω 44 If you think reward-seekers are plausible, you should also think “fitness-seekers” are plausible. But their risks aren't the same. The AI safety community often emphasizes reward-seeking as a central case of a misaligned AI alongside scheming (e.g., Cotra’s sycophant vs schemer, Carlsmith’s terminal vs instrumental traini
Explore this link on the map →saved by
related reading
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Reward is not the optimization target — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- What failure looks like — LessWronglesswrong.com
- Will reward-seekers respond to distant incentives? — LessWronglesswrong.com
- Models Don't "Get Reward" — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- A Toy Environment For Exploring Reasoning About Reward — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Intrinsic Power-Seeking: AI Might Seek Power for Power’s Saketurntrout.com
- Why AIs aren't power-seeking yet — LessWronglesswrong.com