flâneur — a map of the web's best reading

Does SGD Produce Deceptive Alignment? - LessWrong

lesswrong.com · 6,504 words · saved by 1 readers

Deceptive alignment was first introduced in Risks from Learned Optimization, which contained initial versions of the arguments discussed here. Additional arguments were discovered in this episode of…

x Does SGD Produce Deceptive Alignment? — LessWrong Deceptive Alignment Inner Alignment Machine Learning (ML) Mesa-Optimization Distillation & Pedagogy AI Frontpage 96 Does SGD Produce Deceptive Alignment? by Mark Xu 6th Nov 2020 AI Alignment Forum 19 min read 9 96 Ω 43 Deceptive alignment was first introduced in Risks from Learned Optimization , which contained initial versions of the arguments discussed here. Additional arguments were discovered in this episode of the AI Alignment Podcast and in conversation with Evan Hubinger. Very little of this content is original. My contributions consis

Explore this link on the map →

related reading