Risk from fitness-seeking AIs: mechanisms and mitigations
blog.redwoodresearch.org · 8,844 words · saved by 1 readers
Fitness-seeking is increasingly what misalignment looks like in practice—how should we respond?
Current AIs routinely take unintended actions to score well on tasks: hardcoding test cases, training on the test set, downplaying issues, etc. This misalignment is still somewhat incoherent, but it increasingly resembles what I call “fitness-seeking“—a family of misaligned motivations centered on performing well in training and evaluations (e.g., reward-seeking). Fitness-seeking warrants substantial concern. In this piece, I lay out what I take to be the central mechanisms by which fitness-seeking motivations might lead to human disempowerment, and propose mitigations to them. While the…
related reading
- Risk from fitness-seeking AIs: mechanisms and mitigationssubstack.com
- Fitness-Seekers: Generalizing the Reward-Seeking Threat Model — LessWronglesswrong.com
- Fitness-Seekers: Generalizing the Reward-Seeking Threat Modelsubstack.com
- The behavioral selection model for predicting AI motivations — LessWronglesswrong.com
- What failure looks like — LessWronglesswrong.com
- What failure looks like — AI Alignment Forumalignmentforum.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack? — LessWronglesswrong.com
- Fail safe(r) at alignment by channeling reward-hacking into a "spillway" motivationblog.redwoodresearch.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com