Does SGD Produce Deceptive Alignment? - LessWrong
Deceptive alignment was first introduced in Risks from Learned Optimization, which contained initial versions of the arguments discussed here. Additional arguments were discovered in this episode of…
x Does SGD Produce Deceptive Alignment? — LessWrong Deceptive Alignment Inner Alignment Machine Learning (ML) Mesa-Optimization Distillation & Pedagogy AI Frontpage 96 Does SGD Produce Deceptive Alignment? by Mark Xu 6th Nov 2020 AI Alignment Forum 19 min read 9 96 Ω 43 Deceptive alignment was first introduced in Risks from Learned Optimization , which contained initial versions of the arguments discussed here. Additional arguments were discovered in this episode of the AI Alignment Podcast and in conversation with Evan Hubinger. Very little of this content is original. My contributions consis
Explore this link on the map →related reading
- How likely is deceptive alignment? — AI Alignment Forumalignmentforum.org
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Did Claude 3 Opus align itself via gradient hacking? — LessWronglesswrong.com
- Inner Alignment: Explain like I'm 12 Edition — LessWronglesswrong.com
- Deceptive Alignment — AI Alignment Forumalignmentforum.org
- Why AI alignment could be hard with modern deep learningcold-takes.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Relaxed adversarial training for inner alignment — AI Alignment Forumalignmentforum.org