The self-unalignment problem — LessWrong
lesswrong.com · 6,295 words · saved by 1 readers
The usual basic framing of alignment looks something like this: • …
x The self-unalignment problem — LessWrong Coherent Extrapolated Volition Failure mode Subagents Value Learning AI Frontpage 160 The self-unalignment problem by Jan_Kulveit , rosehadshar 14th Apr 2023 AI Alignment Forum 12 min read 24 160 Ω 51 The usual basic framing of alignment looks something like this: We have a system “A” which we are trying to align with system "H", which should establish some alignment relation “f” between the systems. Generally, as the result, the aligned system A should do "what the system H wants". Two things stand out in this basic framing: Alignment is a relation,
related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- ALIGNMENT - by vincent huang - a slice of my mindmindslice.substack.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- The case against AI alignment — LessWronglesswrong.com
- The Artificiality of Alignmentjoinreboot.org
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- What is AI alignment? - by Adam Jones - BlueDot Impactblog.bluedot.org
- What Is The Alignment Problem? — LessWronglesswrong.com
- Explore: AI alignmentarbital.greaterwrong.com
- Value Alignment Is a Pseudo Conceptzilanqian.substack.com
- Another (outer) alignment failure story — AI Alignment Forumalignmentforum.org
- Self-Other Overlap: A Neglected Approach to AI Alignment — LessWronglesswrong.com