flâneur — a map of the web's best reading

Irrationality as a Defense Mechanism for Reward-hacking — LessWrong

lesswrong.com · 2,172 words · saved by 1 readers

This post was written as part of research done at MATS 9.0 under the mentorship of Richard Ngo. It's related to my previous post, but should be reada…

x Irrationality as a Defense Mechanism for Reward-hacking — LessWrong Agent Foundations MATS Program AI World Modeling Frontpage 49 Irrationality as a Defense Mechanism for Reward-hacking by Ashe Vazquez Nuñez 18th Jan 2026 4 min read 8 49 This post was written as part of research done at MATS 9.0 under the mentorship of Richard Ngo. It's related to my previous post , but should be readable as a standalone . Remark: I'm not yet familiar enough with the active inference literature to be sure that the issues I bring up haven't been addressed or discussed. If you think my characterisation of the

Explore this link on the map →

saved by

related reading