flâneur — a map of the web's best reading

The Natural Abstraction Hypothesis: Implications and Evidence — LessWrong

lesswrong.com · 9,067 words · saved by 1 readers

This post was written under Evan Hubinger’s direct guidance and mentorship, as a part of the Stanford Existential Risks Institute ML Alignment Theory Scholars (MATS) program. Additional thanks to Jamie Bernardi, Oly Sourbut and Shawn Hu, for their thoughts and feedback on this post. The Natural Abstraction Hypothesis, proposed by John Wentworth, states that there exist abstractions (relatively low-dimensional summaries which capture information relevant for prediction) which are "natural" in the sense that we should expect a wide variety of cognitive systems to converge on using them. If this hypothesis (which I will refer to as NAH) is true in a very strong sense, then it might be the case that we get "alignment by default", where an unsupervised learner finds a simple embedding of human values, and a supervised learner can be trained with access to this learned model. However, even if the NAH is true in a weaker sense, this has important implications, for areas such as: Several diffe

x The Natural Abstraction Hypothesis: Implications and Evidence — LessWrong ML Alignment Theory Scholars Program Winter 2021 Natural Abstraction Has Diagram Abstraction Interpretability (ML & AI) MATS Program AI Frontpage 44 The Natural Abstraction Hypothesis: Implications and Evidence by CallumMcDougall 14th Dec 2021 AI Alignment Forum 23 min read 9 44 Ω 24 This post was written under Evan Hubinger’s direct guidance and mentorship, as a part of the Stanford Existential Risks Institute ML Alignment Theory Scholars (MATS) program . Additional thanks to Jamie Bernardi, Oly Sourbut and Shawn Hu,

Explore this link on the map →

related reading