The Computational Anatomy of Human Values - LessWrong
Tldr: Human values are primarily linguistic concepts encoded via webs of association and valence in the cortex learnt through unsupervised (primarily linguistic) learning. These value concepts are bo…
x The Computational Anatomy of Human Values — LessWrong Human Values Outer Alignment Value Learning AI Rationality World Modeling Frontpage 76 The Computational Anatomy of Human Values by beren 6th Apr 2023 AI Alignment Forum 36 min read 30 76 Ω 29 This is crossposted from my personal blog . Epistemic Status : Much of this draws from my studies in neuroscience and ML. Many of the ideas in this post are heavily inspired by the work of Steven Byrnes and the authors of Shard Theory . However, it speculates quite a long way in advance of the scientific frontier and is almost certainly incorrect in
Explore this link on the map →related reading
- The Shard Theory of Human Valuesturntrout.com
- The shard theory of human values — LessWronglesswrong.com
- A Crash Course in the Neuroscience of Human Motivation — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- The shard theory of human values — AI Alignment Forumalignmentforum.org
- Reward is not the optimization target — LessWronglesswrong.com
- Andrej Karpathy — AGI is still a decade awaydwarkesh.com
- Alignment By Default — AI Alignment Forumalignmentforum.org
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- Value Learning Needs a Low-Dimensional Bottleneck — LessWronglesswrong.com
- Value systematization: how values become coherent (and misaligned) — AI Alignment Forumalignmentforum.org
- Models Don't "Get Reward" — LessWronglesswrong.com