How likely is deceptive alignment? — AI Alignment Forum
The following is an edited transcript of a talk I gave. I have given this talk at multiple places, including first at Anthropic and then for ELK winners and at Redwood Research, though the version that this document is based on is the version I gave to SERI MATS fellows. Thanks to Jonathan Ng, Ryan Kidd, and others for help transcribing that talk. Substantial edits were done on top of the transcription by me. Though all slides are embedded below, the full slide deck is also available here. Today I’m going to be talking about deceptive alignment. Deceptive alignment is something I'm very concerned about and is where I think most of the existential risk from AI comes from. And I'm going to try to make the case for why I think that this is the default outcome of machine learning. First of all, what am I talking about? I want to disambiguate between two closely related, but distinct concepts. The first concept is dishonesty. This is something that many people are concerned about in models,
x How likely is deceptive alignment? — AI Alignment Forum Deceptive Alignment Deception Inner Alignment AI Frontpage 50 How likely is deceptive alignment? by evhub 30th Aug 2022 72 min read 31 50 The following is an edited transcript of a talk I gave. I have given this talk at multiple places, including first at Anthropic and then for ELK winners and at Redwood Research, though the version that this document is based on is the version I gave to SERI MATS fellows. Thanks to Jonathan Ng, Ryan Kidd, and others for help transcribing that talk. Substantial edits were done on top of the transcriptio
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Does SGD Produce Deceptive Alignment? — LessWronglesswrong.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Deep Deceptiveness — LessWronglesswrong.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Why AI alignment could be hard with modern deep learningcold-takes.com