flâneur — a map of the web's best reading

Value systematization: how values become coherent (and misaligned) — AI Alignment Forum

alignmentforum.org · 14,708 words · saved by 1 readers

Many discussions of AI risk are unproductive or confused because it’s hard to pin down concepts like “coherence” and “expected utility maximization” in the context of deep learning. In this post I attempt to bridge this gap by describing a process by which AI values might become more coherent, which I’m calling “value systematization”, and which plays a crucial role in my thinking about AI risk. I define value systematization as the process of an agent learning to represent its previous values as examples or special cases of other simpler and more broadly-scoped values. I think of value systematization as the most plausible mechanism by which AGIs might acquire broadly-scoped misaligned goals which incentivize takeover. I’ll first discuss the related concept of belief systematization. I’ll next characterize what value systematization looks like in humans, to provide some intuitions. I’ll then talk about what value systematization might look like in AIs. I think of value systematization

x Value systematization: how values become coherent (and misaligned) — AI Alignment Forum Understanding systematization Value Learning AI Frontpage 46 Value systematization: how values become coherent (and misaligned) by Richard_Ngo 27th Oct 2023 16 min read 49 46 Many discussions of AI risk are unproductive or confused because it’s hard to pin down concepts like “coherence” and “expected utility maximization” in the context of deep learning. In this post I attempt to bridge this gap by describing a process by which AI values might become more coherent, which I’m calling “value systematization

Explore this link on the map →

related reading