flâneur

Risk from fitness-seeking AIs: mechanisms and mitigations

substack.com · 8,844 words · saved by 1 readers

Fitness-seeking is increasingly what misalignment looks like in practice—how should we respond?

Current AIs routinely take unintended actions to score well on tasks: hardcoding test cases, training on the test set, downplaying issues, etc. This misalignment is still somewhat incoherent, but it increasingly resembles what I call “fitness-seeking“—a family of misaligned motivations centered on performing well in training and evaluations (e.g., reward-seeking). Fitness-seeking warrants substantial concern. In this piece, I lay out what I take to be the central mechanisms by which fitness-seeking motivations might lead to human disempowerment, and propose mitigations to them. While the…

saved by

related reading