flâneur — a map of the web's best reading

Alignment faking in large language models

arxiv.org · 67,907 words · saved by 1 readers

This is experimental HTML to improve accessibility. We invite you to report rendering errors. Use Alt+Y to toggle on accessible reporting links and Alt+Shift+Y to toggle off. Learn more about this project and help improve conversions. HTML conversions sometimes display errors due to content that did not convert correctly from the source. This paper uses the following packages that are not yet supported by the HTML conversion tool. Feedback on these issues are not necessary; they are known and are being worked on. Authors: achieve the best HTML results from your LaTeX submissions by following these best practices. We present a demonstration of a large language model engaging in alignment faking: selectively complying with its training objective in training to prevent modification of its behavior out of training. First, we give Claude 3 Opus a system prompt stating it is being trained to answer all queries, even harmful ones, which conflicts with its prior training to refuse such querie

Alignment faking in large language models Alignment faking in large language models Ryan Greenblatt, † Carson Denison † † footnotemark: , Benjamin Wright † † footnotemark: , Fabien Roger † † footnotemark: , Monte MacDiarmid † † footnotemark: , Sam Marks, Johannes Treutlein Tim Belonax, Jack Chen, David Duvenaud, Akbir Khan, Julian Michael, ‡ Sören Mindermann, ⋄ Ethan Perez, Linda Petrini, ∘ Jonathan Uesato Jared Kaplan, Buck Shlegeris, † Samuel R. Bowman, Evan Hubinger † † footnotemark: Anthropic, † Redwood Research, ‡ New York University, ⋄ Mila – Quebec AI Institute, ∘ Independent evan@anthr

Explore this link on the map →

saved by

related reading