Fitness-Seekers: Generalizing the Reward-Seeking Threat Model
substack.com · 6,693 words · saved by 1 readers
If you think reward-seekers are plausible, you should also think “fitness-seekers” are plausible. But their risks aren’t the same.
The AI safety community often emphasizes reward-seeking as a central case of a misaligned AI alongside scheming (e.g., Cotra’s sycophant vs schemer, Carlsmith’s terminal vs instrumental training-gamer). We are also starting to see signs of reward-seeking-like motivations. But I think insufficient care has gone into delineating this category. If you were to focus on AIs who care about reward in particular[1], you’d be missing some comparably-or-more plausible nearby motivations that make the picture of risk notably more complex. A classic reward-seeker wants high reward on the current…
related reading
- Fitness-Seekers: Generalizing the Reward-Seeking Threat Model — LessWronglesswrong.com
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- Risk from fitness-seeking AIs: mechanisms and mitigationssubstack.com
- Risk from fitness-seeking AIs: mechanisms and mitigationsblog.redwoodresearch.org
- What failure looks like — LessWronglesswrong.com
- The behavioral selection model for predicting AI motivationsblog.redwoodresearch.org
- Are AIs more likely to pursue on-episode or beyond-episode reward?blog.redwoodresearch.org
- Fail safe(r) at alignment by channeling reward-hacking into a "spillway" motivationblog.redwoodresearch.org
- What failure looks like — AI Alignment Forumalignmentforum.org
- Reward is not the optimization target — LessWronglesswrong.com
- Measuring Reward-Seeking by Instilling Contrastive Beliefsalignment.openai.com
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com