flâneur

Fitness-Seekers: Generalizing the Reward-Seeking Threat Model

substack.com · 6,693 words · saved by 1 readers

If you think reward-seekers are plausible, you should also think “fitness-seekers” are plausible. But their risks aren’t the same.

The AI safety community often emphasizes reward-seeking as a central case of a misaligned AI alongside scheming (e.g., Cotra’s sycophant vs schemer, Carlsmith’s terminal vs instrumental training-gamer). We are also starting to see signs of reward-seeking-like motivations. But I think insufficient care has gone into delineating this category. If you were to focus on AIs who care about reward in particular[1], you’d be missing some comparably-or-more plausible nearby motivations that make the picture of risk notably more complex. A classic reward-seeker wants high reward on the current…

related reading