flâneur — a map of the web's best reading

A Toy Environment For Exploring Reasoning About Reward — LessWrong

lesswrong.com · 3,361 words · saved by 1 readers

tldr: We share a toy environment that we found useful for understanding how reasoning changed over the course of capabilities-focused RL. Over the co…

x A Toy Environment For Exploring Reasoning About Reward — LessWrong AI Frontpage 56 A Toy Environment For Exploring Reasoning About Reward by jenny , Bronson Schoen 25th Mar 2026 AI Alignment Forum 3 min read 7 56 Ω 26 tldr : We share a toy environment that we found useful for understanding how reasoning changed over the course of capabilities-focused RL . Over the course of capabilities-focused RL, the model biases more strongly towards reward hints over direct instruction in this environment. Setup When we noticed the increase in verbalized alignment evaluation awareness during capabilities

Explore this link on the map →

saved by

related reading