The Natural Abstraction Hypothesis: Implications and Evidence — LessWrong
This post was written under Evan Hubinger’s direct guidance and mentorship, as a part of the Stanford Existential Risks Institute ML Alignment Theory Scholars (MATS) program. Additional thanks to Jamie Bernardi, Oly Sourbut and Shawn Hu, for their thoughts and feedback on this post. The Natural Abstraction Hypothesis, proposed by John Wentworth, states that there exist abstractions (relatively low-dimensional summaries which capture information relevant for prediction) which are "natural" in the sense that we should expect a wide variety of cognitive systems to converge on using them. If this hypothesis (which I will refer to as NAH) is true in a very strong sense, then it might be the case that we get "alignment by default", where an unsupervised learner finds a simple embedding of human values, and a supervised learner can be trained with access to this learned model. However, even if the NAH is true in a weaker sense, this has important implications, for areas such as: Several diffe
x The Natural Abstraction Hypothesis: Implications and Evidence — LessWrong ML Alignment Theory Scholars Program Winter 2021 Natural Abstraction Has Diagram Abstraction Interpretability (ML & AI) MATS Program AI Frontpage 44 The Natural Abstraction Hypothesis: Implications and Evidence by CallumMcDougall 14th Dec 2021 AI Alignment Forum 23 min read 9 44 Ω 24 This post was written under Evan Hubinger’s direct guidance and mentorship, as a part of the Stanford Existential Risks Institute ML Alignment Theory Scholars (MATS) program . Additional thanks to Jamie Bernardi, Oly Sourbut and Shawn Hu,
Explore this link on the map →related reading
- Dialogue: Is there a Natural Abstraction of Good? — LessWronglesswrong.com
- Natural Abstractions: Key Claims, Theorems, and Critiques — LessWronglesswrong.com
- 2023 letter | Zhengdongzhengdongwang.com
- Natural Language Autoencoders \ Anthropicanthropic.com
- Faithful, Interpretable Model Explanations via Causal Abstraction | SAIL Blogai.stanford.edu
- The Plan - 2023 Version — LessWronglesswrong.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- What If We Had Bigger Brains? Imagining Minds beyond Ours-Stephen Wolfram Writingswritings.stephenwolfram.com
- The Building Blocks of Interpretabilitydistill.pub
- A Pragmatic Vision for Interpretability — AI Alignment Forumalignmentforum.org
- The "Minimal Latents" Approach to Natural Abstractions — AI Alignment Forumalignmentforum.org
- Interpretability Dreamstransformer-circuits.pub