Why do models task game? - LessWrong 2.0 viewer
We see this as a work of high-level model forensics. Rather than investigating a single incident, the core problem here is taking an ambiguous pattern of behavior across many contexts with various plausible motivations, and practicing how to distinguish the motivations. The main subject of study is DeepSeek v4 Pro, but we also report results on other models. We obtain several of our key results from a realistic long-horizon coding environment that induces task gaming. We open-source all our environments here.
TL;DR How can we study misalignment with today’s models as proxies? They’re clearly not paperclip maximizers, but they also often do things the user doesn’t want. A strong contender for a real misaligned propensity is task gaming: taking actions that don’t complete a task but superficially seem like they do, such as hardcoding tests or falsely claiming a task is fully complete. But maybe task gaming is just a crude heuristic, or the model mistakenly trying to achieve the user’s intent? In this post we do a deep dive into why a range of models task game. We see this as a work of high-level…
saved by
related reading
- [2606.26071] Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignmentarxiv.org
- [2607.18966] Measuring Reward-Seeking via Contrastive Belief Updatesarxiv.org
- Measuring Reward-Seeking by Instilling Contrastive Beliefsalignment.openai.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- The Case for Model Forensics — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- confessions_paper.pdfcdn.openai.com
- Why We Are Excited About Confessionsalignment.openai.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com