flâneur — a map of the web's best reading

The self-unalignment problem — LessWrong

lesswrong.com · 6,295 words · saved by 1 readers

The usual basic framing of alignment looks something like this: • …

x The self-unalignment problem — LessWrong Coherent Extrapolated Volition Failure mode Subagents Value Learning AI Frontpage 160 The self-unalignment problem by Jan_Kulveit , rosehadshar 14th Apr 2023 AI Alignment Forum 12 min read 24 160 Ω 51 The usual basic framing of alignment looks something like this: We have a system “A” which we are trying to align with system "H", which should establish some alignment relation “f” between the systems. Generally, as the result, the aligned system A should do "what the system H wants". Two things stand out in this basic framing: Alignment is a relation,

Explore this link on the map →

related reading