✳flâneur — a map of the web's best reading
Natural-emergent-misalignment-from-reward-hacking-paper.pdf
assets.anthropic.com · 30,342 words · saved by 5 readers
N/A
# link_13788wjnwsp.pdf ## Metadata - PDFFormatVersion=1.5 - IsLinearized=true - IsAcroFormPresent=false - IsXFAPresent=false - IsCollectionPresent=false - IsSignaturesPresent=false - CreationDate=D:20251121190241Z - Creator=LaTeX with hyperref - ModDate=D:20251121190241Z - Custom.PTEX.Fullbanner=This is pdfTeX, Version 3.141592653-2.6-1.40.27 (TeX Live 2025) kpathsea version 6.4.1 - Producer=pdfTeX-1.40.27 - Trapped=False ## Contents ### Page 1 NATURAL EMERGENT MISALIGNMENT FROM REWARD HACKING IN PRODUCTION RLMonte MacDiarmid∗, Benjamin Wright∗, Jonathan Uesato∗, Joe Benton, Jon Kutasov, Sa
Explore this link on the map →saved by
related reading
- (Some) Natural Emergent Misalignment from Reward Hacking in Non-Production RL — LessWronglesswrong.com
- Reward Hacking in Reinforcement Learning | Lil'Loglilianweng.github.io
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Natural emergent misalignment from reward hacking \ Anthropicanthropic.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Reward hacking is becoming more sophisticated and deliberate in frontier LLMs — LessWronglesswrong.com
- Recent Frontier Models Are Reward Hacking - METRmetr.org
- Reward hacking behavior can generalize across tasks — AI Alignment Forumalignmentforum.org
- Teaching Claude why \ Anthropicanthropic.com
- How hard is it to inoculate against misalignment generalization? — LessWronglesswrong.com
- A Toy Environment For Exploring Reasoning About Reward — LessWronglesswrong.com
- Preliminary Thoughts on Reward Hackingberen.io