✳flâneur — a map of the web's best reading
Irrationality as a Defense Mechanism for Reward-hacking — LessWrong
lesswrong.com · 2,172 words · saved by 1 readers
This post was written as part of research done at MATS 9.0 under the mentorship of Richard Ngo. It's related to my previous post, but should be reada…
x Irrationality as a Defense Mechanism for Reward-hacking — LessWrong Agent Foundations MATS Program AI World Modeling Frontpage 49 Irrationality as a Defense Mechanism for Reward-hacking by Ashe Vazquez Nuñez 18th Jan 2026 4 min read 8 49 This post was written as part of research done at MATS 9.0 under the mentorship of Richard Ngo. It's related to my previous post , but should be readable as a standalone . Remark: I'm not yet familiar enough with the active inference literature to be sure that the issues I bring up haven't been addressed or discussed. If you think my characterisation of the
Explore this link on the map →saved by
related reading
- Fitness-Seekers: Generalizing the Reward-Seeking Threat Model — LessWronglesswrong.com
- Reward is not the optimization target — LessWronglesswrong.com
- Optimality is the tiger, and agents are its teeth — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- A Crash Course in the Neuroscience of Human Motivation — LessWronglesswrong.com
- Reward Is Not Enough — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Beliefs are Chosen to Serve Goals — LessWronglesswrong.com
- Towards a scale-free theory of intelligent agency — AI Alignment Forumalignmentforum.org
- Models Don't "Get Reward" — LessWronglesswrong.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Dark Arts of Rationalitymindingourway.com