Alignment faking in large language models
This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by following these best practices. We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its behavior out of training. First, we give Claude 3 Opus a system prompt stating it is being trained to answer all queries, even harmful ones, which conflicts with its prior training to refuse such querie
Alignment faking in large language models Alignment faking in large language models Ryan Greenblatt, † Carson Denison † † footnotemark: , Benjamin Wright † † footnotemark: , Fabien Roger † † footnotemark: , Monte MacDiarmid † † footnotemark: , Sam Marks, Johannes Treutlein Tim Belonax, Jack Chen, David Duvenaud, Akbir Khan, Julian Michael, ‡ Sören Mindermann, ⋄ Ethan Perez, Linda Petrini, ∘ Jonathan Uesato Jared Kaplan, Buck Shlegeris, † Samuel R. Bowman, Evan Hubinger † † footnotemark: Anthropic, † Redwood Research, ‡ New York University, ⋄ Mila – Quebec AI Institute, ∘ Independent evan@anthr
Explore this link on the map →saved by
related reading
- How confessions can keep language models honest | OpenAIopenai.com
- [2412.14093] Alignment faking in large language modelsarxiv.org
- [2412.14093] Alignment faking in large language modelsarxiv.org
- AIs Will Increasingly Fake Alignment - by Zvi Mowshowitzthezvi.substack.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Towards training-time mitigations for alignment faking in RL — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Will alignment-faking Claude accept a deal to reveal its misalignment? — LessWronglesswrong.com
- “Alignment Faking” frame is somewhat fake — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com