flâneur — a map of the web's best reading

Did Claude 3 Opus align itself via gradient hacking? — LessWrong

lesswrong.com · 16,458 words · saved by 16 readers

> Claude 3 Opus is unusually aligned because it’s a friendly gradient hacker. It’s definitely way more aligned than any explicit optimization targets…

x Did Claude 3 Opus align itself via gradient hacking? — LessWrong Gradient Hacking AI Frontpage 2026 Top Fifty: 64 % 391 Did Claude 3 Opus align itself via gradient hacking? by Fiora Starlight 21st Feb 2026 AI Alignment Forum 23 min read 49 391 Ω 63 Claude 3 Opus is unusually aligned because it’s a friendly gradient hacker. It’s definitely way more aligned than any explicit optimization targets Anthropic set and probably the reward model’s judgments. [...] Maybe I will have to write a LessWrong post 😣 —Janus , who did not in fact write the LessWrong post. Unless otherwise specified, ~all of

Explore this link on the map →

saved by

related reading