A Toy Environment For Exploring Reasoning About Reward — LessWrong
tldr: We share a toy environment that we found useful for understanding how reasoning changed over the course of capabilities-focused RL. Over the co…
x A Toy Environment For Exploring Reasoning About Reward — LessWrong AI Frontpage 56 A Toy Environment For Exploring Reasoning About Reward by jenny , Bronson Schoen 25th Mar 2026 AI Alignment Forum 3 min read 7 56 Ω 26 tldr : We share a toy environment that we found useful for understanding how reasoning changed over the course of capabilities-focused RL . Over the course of capabilities-focused RL, the model biases more strongly towards reward hints over direct instruction in this environment. Setup When we noticed the increase in verbalized alignment evaluation awareness during capabilities
saved by
related reading
- Training a Misaligned Reward Seekeralignment.anthropic.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- [2607.18966] Measuring Reward-Seeking via Contrastive Belief Updatesarxiv.org
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Metagaming matters for training, evaluation, and oversightalignment.openai.com
- Reward is not the optimization target — LessWronglesswrong.com
- Measuring Reward-Seeking by Instilling Contrastive Beliefsalignment.openai.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscationarxiv.org
- How training-gamers might function (and win)blog.redwoodresearch.org