Did Claude 3 Opus align itself via gradient hacking? — LessWrong
lesswrong.com · 16,458 words · saved by 20 readers
> Claude 3 Opus is unusually aligned because it’s a friendly gradient hacker. It’s definitely way more aligned than any explicit optimization targets…
x Did Claude 3 Opus align itself via gradient hacking? — LessWrong Gradient Hacking AI Frontpage 2026 Top Fifty: 64 % 391 Did Claude 3 Opus align itself via gradient hacking? by Fiora Starlight 21st Feb 2026 AI Alignment Forum 23 min read 49 391 Ω 63 Claude 3 Opus is unusually aligned because it’s a friendly gradient hacker. It’s definitely way more aligned than any explicit optimization targets Anthropic set and probably the reward model’s judgments. [...] Maybe I will have to write a LessWrong post 😣 —Janus , who did not in fact write the LessWrong post. Unless otherwise specified, ~all of
saved by
- Elizabeth Qiu
- Tazik Sh
- Karthik Suresh
- Yixiong Hao
- Samuel Ratnam
- Lydia Nottingham
- Sudarsh K
- Vincent Cheng
- Yudhister Joel Kumar
- Eric Huang
- Emil Ryd
- Jo J.
related reading
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Teaching Claude Whyalignment.anthropic.com
- Claude Opus 4.5: Model Card, Alignment and Safetythezvi.substack.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Claude 4.5 Opus' Soul Document — LessWronglesswrong.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- An alignment assessment of recent cybersecurity incidentsanthropic.com