Irrationality as a Defense Mechanism for Reward-hacking — LessWrong
lesswrong.com · 2,172 words · saved by 1 readers
This post was written as part of research done at MATS 9.0 under the mentorship of Richard Ngo. It's related to my previous post, but should be reada…
x Irrationality as a Defense Mechanism for Reward-hacking — LessWrong Agent Foundations MATS Program AI World Modeling Frontpage 49 Irrationality as a Defense Mechanism for Reward-hacking by Ashe Vazquez Nuñez 18th Jan 2026 4 min read 8 49 This post was written as part of research done at MATS 9.0 under the mentorship of Richard Ngo. It's related to my previous post , but should be readable as a standalone . Remark: I'm not yet familiar enough with the active inference literature to be sure that the issues I bring up haven't been addressed or discussed. If you think my characterisation of the
saved by
related reading
- Fitness-Seekers: Generalizing the Reward-Seeking Threat Model — LessWronglesswrong.com
- Reward is not the optimization target — LessWronglesswrong.com
- Optimality is the tiger, and agents are its teeth — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Dark Arts of Rationality — LessWronglesswrong.com
- Towards a scale-free theory of intelligent agency — AI Alignment Forumalignmentforum.org
- Towards a scale-free theory of intelligent agencymindthefuture.info
- Embedded Agency (full-text version) — LessWronglesswrong.com
- Reward Is Not the Optimization Targetturntrout.com
- Reward Is Not Enough — LessWronglesswrong.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Beliefs are Chosen to Serve Goals — LessWronglesswrong.com