Specification gaming examples in AI — LessWrong
Various examples (and lists of examples) of unintended behaviors in AI systems have appeared in recent years. One interesting type of unintended behavior is finding a way to game the specified objective: generating a solution that literally satisfies the stated objective but fails to solve the problem according to the human designer’s intent. This occurs when the objective is poorly specified, and includes reinforcement learning agents hacking the reward function, evolutionary algorithms gaming the fitness function, etc. While ‘specification gaming’ is a somewhat vague category, it is particularly referring to behaviors that are clearly hacks, not just suboptimal solutions. Since these examples are currently scattered across several lists, I have put together a master list of examples collected from the various existing sources. This list is intended to be comprehensive and up-to-date, and serve as a resource for AI safety research and discussion. If you know of any interesting example
x Specification gaming examples in AI — LessWrong Best of LessWrong 2018 Goodhart's Law AI Risk AI Frontpage 49 Specification gaming examples in AI by Vika 3rd Apr 2018 AI Alignment Forum 1 min read 9 49 Ω 18 [Cross-posted from personal blog ] Various examples (and lists of examples ) of unintended behaviors in AI systems have appeared in recent years. One interesting type of unintended behavior is finding a way to game the specified objective: generating a solution that literally satisfies the stated objective but fails to solve the problem according to the human designer’s intent. This occur
Explore this link on the map →saved by
related reading
- Specification gaming: the flip side of AI ingenuity — Google DeepMinddeepmind.google
- Specification gaming: the flip side of AI ingenuity — Google DeepMinddeepmind.google
- Specification gaming: the flip side of AI ingenuity — Google DeepMinddeepmind.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Mediumdeepmindsafetyresearch.medium.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- How Go Players Disempower Themselves to AI — LessWronglesswrong.com
- Recontextualization Mitigates Specification Gaming Without Modifying the Specification — LessWronglesswrong.com
- [2602.12316] GT-HarmBench: Benchmarking AI Safety Risks Through the Lens of Game Theoryarxiv.org
- AI Safety Seems Hard to Measurecold-takes.com