Does SGD Produce Deceptive Alignment? - LessWrong
Deceptive alignment was first introduced in Risks from Learned Optimization, which contained initial versions of the arguments discussed here. Additional arguments were discovered in this episode of…
x Does SGD Produce Deceptive Alignment? — LessWrong Deceptive Alignment Inner Alignment Machine Learning (ML) Mesa-Optimization Distillation & Pedagogy AI Frontpage 96 Does SGD Produce Deceptive Alignment? by Mark Xu 6th Nov 2020 AI Alignment Forum 19 min read 9 96 Ω 43 Deceptive alignment was first introduced in Risks from Learned Optimization , which contained initial versions of the arguments discussed here. Additional arguments were discovered in this episode of the AI Alignment Podcast and in conversation with Evan Hubinger. Very little of this content is original. My contributions consis
related reading
- How likely is deceptive alignment? — AI Alignment Forumalignmentforum.org
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Alignment is not solved but it increasingly looks solvablealigned.substack.com
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Deceptive Alignment — AI Alignment Forumalignmentforum.org
- Alignment Faking Mitigationsalignment.anthropic.com
- Self-Other Overlap: A Neglected Approach to AI Alignment — LessWronglesswrong.com