Value systematization: how values become coherent (and misaligned) — AI Alignment Forum
Many discussions of AI risk are unproductive or confused because it’s hard to pin down concepts like “coherence” and “expected utility maximization” in the context of deep learning. In this post I attempt to bridge this gap by describing a process by which AI values might become more coherent, which I’m calling “value systematization”, and which plays a crucial role in my thinking about AI risk. I define value systematization as the process of an agent learning to represent its previous values as examples or special cases of other simpler and more broadly-scoped values. I think of value systematization as the most plausible mechanism by which AGIs might acquire broadly-scoped misaligned goals which incentivize takeover. I’ll first discuss the related concept of belief systematization. I’ll next characterize what value systematization looks like in humans, to provide some intuitions. I’ll then talk about what value systematization might look like in AIs. I think of value systematization
x Value systematization: how values become coherent (and misaligned) — AI Alignment Forum Understanding systematization Value Learning AI Frontpage 46 Value systematization: how values become coherent (and misaligned) by Richard_Ngo 27th Oct 2023 16 min read 49 46 Many discussions of AI risk are unproductive or confused because it’s hard to pin down concepts like “coherence” and “expected utility maximization” in the context of deep learning. In this post I attempt to bridge this gap by describing a process by which AI values might become more coherent, which I’m calling “value systematization
Explore this link on the map →related reading
- Value systematization: how values become coherent (and misaligned) — LessWronglesswrong.com
- The hot mess theory of AI misalignment: More intelligent agents behave less coherently | Jascha’s blogsohl-dickstein.github.io
- Alignment By Default — AI Alignment Forumalignmentforum.org
- The shard theory of human values — LessWronglesswrong.com
- ⿻ Symbiogenesis vs. Convergent Consequentialism — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- Why AI alignment could be hard with modern deep learningcold-takes.com
- Complexity of value — LessWronglesswrong.com
- What’s Your P(WEIRD)? — LessWronglesswrong.com
- What Is The Alignment Problem? — LessWronglesswrong.com
- LLM Alignment, ethical and mathematical realism, and the most important actions in davidad's understanding — LessWronglesswrong.com