Value systematization: how values become coherent (and misaligned) — LessWrong
Many discussions of AI risk are unproductive or confused because it’s hard to pin down concepts like “coherence” and “expected utility maximization” in the context of deep learning. In this post I attempt to bridge this gap by describing a process by which AI values might become more coherent, which I’m calling “value systematization”, and which plays a crucial role in my thinking about AI risk. I define value systematization as the process of an agent learning to represent its previous values as examples or special cases of other simpler and more broadly-scoped values. I think of value systematization as the most plausible mechanism by which AGIs might acquire broadly-scoped misaligned goals which incentivize takeover. I’ll first discuss the related concept of belief systematization. I’ll next characterize what value systematization looks like in humans, to provide some intuitions. I’ll then talk about what value systematization might look like in AIs. I think of value systematization
x Value systematization: how values become coherent (and misaligned) — LessWrong Understanding systematization Value Learning AI Frontpage 112 Value systematization: how values become coherent (and misaligned) by Richard_Ngo 27th Oct 2023 AI Alignment Forum 16 min read 49 112 Ω 46 Many discussions of AI risk are unproductive or confused because it’s hard to pin down concepts like “coherence” and “expected utility maximization” in the context of deep learning. In this post I attempt to bridge this gap by describing a process by which AI values might become more coherent, which I’m calling “valu
Explore this link on the map →related reading
- Value systematization: how values become coherent (and misaligned) — AI Alignment Forumalignmentforum.org
- The hot mess theory of AI misalignment: More intelligent agents behave less coherently | Jascha’s blogsohl-dickstein.github.io
- The shard theory of human values — LessWronglesswrong.com
- ⿻ Symbiogenesis vs. Convergent Consequentialism — LessWronglesswrong.com
- Alignment By Default — AI Alignment Forumalignmentforum.org
- What Does It Mean to Align AI With Human Values? | Quanta Magazinequantamagazine.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Complexity of value — LessWronglesswrong.com
- What’s Your P(WEIRD)? — LessWronglesswrong.com
- LLM Alignment, ethical and mathematical realism, and the most important actions in davidad's understanding — LessWronglesswrong.com
- Alignment Is Proven To Be Solvable - by SE Gygesverysane.ai
- 1a3orn's Shortform — LessWronglesswrong.com