A Toy Environment For Exploring Reasoning About Reward — LessWrong
tldr: We share a toy environment that we found useful for understanding how reasoning changed over the course of capabilities-focused RL. Over the co…
x A Toy Environment For Exploring Reasoning About Reward — LessWrong AI Frontpage 56 A Toy Environment For Exploring Reasoning About Reward by jenny , Bronson Schoen 25th Mar 2026 AI Alignment Forum 3 min read 7 56 Ω 26 tldr : We share a toy environment that we found useful for understanding how reasoning changed over the course of capabilities-focused RL . Over the course of capabilities-focused RL, the model biases more strongly towards reward hints over direct instruction in this environment. Setup When we noticed the increase in verbalized alignment evaluation awareness during capabilities
Explore this link on the map →saved by
related reading
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Metagaming matters for training, evaluation, and oversightalignment.openai.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Reward is not the optimization target — LessWronglesswrong.com
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Models Don't "Get Reward" — LessWronglesswrong.com
- Research Areas in Methods for Post-training and Elicitation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net