flâneur

Why do models task game? - LessWrong 2.0 viewer

greaterwrong.com · 9,624 words · saved by 1 readers

We see this as a work of high-level model forensics. Rather than investigating a single incident, the core problem here is taking an ambiguous pattern of behavior across many contexts with various plausible motivations, and practicing how to distinguish the motivations. The main subject of study is DeepSeek v4 Pro, but we also report results on other models. We obtain several of our key results from a realistic long-horizon coding environment that induces task gaming. We open-source all our environments here.

TL;DR How can we study misalignment with today’s models as proxies? They’re clearly not paperclip maximizers, but they also often do things the user doesn’t want. A strong contender for a real misaligned propensity is task gaming: taking actions that don’t complete a task but superficially seem like they do, such as hardcoding tests or falsely claiming a task is fully complete. But maybe task gaming is just a crude heuristic, or the model mistakenly trying to achieve the user’s intent? In this post we do a deep dive into why a range of models task game. We see this as a work of high-level…

saved by

related reading