✳flâneur — a map of the web's best reading
(Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWrong
lesswrong.com · 8,214 words · saved by 1 readers
Authors: Satvik Golechha*, Sid Black*, Joseph Bloom • * Equal Contribution. …
x (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWrong AI Frontpage 2026 Top Fifty: 14 % 127 (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL by 7vik , Sid Black , Joseph Bloom 30th Mar 2026 AI Alignment Forum 21 min read 6 127 Ω 47 Authors: Satvik Golechha*, Sid Black*, Joseph Bloom * Equal Contribution. This work was done as part of the Model Transparency team at the UK AI Security Institute (AISI). Our code is available on GitHub and the model checkpoints and data is available on HuggingFace . Executive Summary In Natural Eme
Explore this link on the map →saved by
related reading
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Natural emergent misalignment from reward hacking \ Anthropicanthropic.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- A Toy Environment For Exploring Reasoning About Reward — LessWronglesswrong.com
- Paper: Prompt Optimization Makes Misalignment Legible — LessWronglesswrong.com
- Training on Documents About Reward Hacking Induces Reward Hacking — LessWronglesswrong.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Research Areas in Evaluation and Guarantees in Reinforcement Learning (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- Training on Documents about Reward Hacking Induces Reward Hackingalignment.anthropic.com