✳flâneur — a map of the web's best reading
Did Claude 3 Opus align itself via gradient hacking? — LessWrong
lesswrong.com · 16,458 words · saved by 16 readers
> Claude 3 Opus is unusually aligned because it’s a friendly gradient hacker. It’s definitely way more aligned than any explicit optimization targets…
x Did Claude 3 Opus align itself via gradient hacking? — LessWrong Gradient Hacking AI Frontpage 2026 Top Fifty: 64 % 391 Did Claude 3 Opus align itself via gradient hacking? by Fiora Starlight 21st Feb 2026 AI Alignment Forum 23 min read 49 391 Ω 63 Claude 3 Opus is unusually aligned because it’s a friendly gradient hacker. It’s definitely way more aligned than any explicit optimization targets Anthropic set and probably the reward model’s judgments. [...] Maybe I will have to write a LessWrong post 😣 —Janus , who did not in fact write the LessWrong post. Unless otherwise specified, ~all of
Explore this link on the map →saved by
- Elizabeth Qiu
- Yixiong Hao
- Samuel Ratnam
- Lydia Nottingham
- Sudarsh K
- Vincent Cheng
- Yudhister Joel Kumar
- Eric Huang
- Emil Ryd
- Jo J.
- Neil Rathi
- Christine Ye
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Automated Alignment Researchers: Using large language models to scale scalable oversight \ Anthropicanthropic.com
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- Claude’s Character \ Anthropicanthropic.com
- Alignment faking in large language modelsarxiv.org
- Opus 4.8 Part 2: Model Welfare - by Zvi Mowshowitzthezvi.substack.com
- Claude 4 System Cardwww-cdn.anthropic.com
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- AIs Will Increasingly Fake Alignment - by Zvi Mowshowitzthezvi.substack.com