flâneur — a map of the web's best reading

Training on Documents About Reward Hacking Induces Reward Hacking — LessWrong

lesswrong.com · 2,402 words · saved by 1 readers

This is a blog post reporting some preliminary work from the Anthropic Alignment Science team, which might be of interest to researchers working acti…

x Training on Documents About Reward Hacking Induces Reward Hacking — LessWrong Self Fulfilling/Refuting Prophecies AI Frontpage 2025 Top Fifty: 5 % 135 Training on Documents About Reward Hacking Induces Reward Hacking by evhub , Nathan Hu 21st Jan 2025 AI Alignment Forum 2 min read 15 135 Ω 61 This is a linkpost for https://alignment.anthropic.com/2025/reward-hacking-ooc/ This is a blog post reporting some preliminary work from the Anthropic Alignment Science team, which might be of interest to researchers working actively in this space. We'd ask you to treat these results like those of a col

Explore this link on the map →

related reading