✳flâneur — a map of the web's best reading
The self-unalignment problem — LessWrong
lesswrong.com · 6,295 words · saved by 1 readers
The usual basic framing of alignment looks something like this: • …
x The self-unalignment problem — LessWrong Coherent Extrapolated Volition Failure mode Subagents Value Learning AI Frontpage 160 The self-unalignment problem by Jan_Kulveit , rosehadshar 14th Apr 2023 AI Alignment Forum 12 min read 24 160 Ω 51 The usual basic framing of alignment looks something like this: We have a system “A” which we are trying to align with system "H", which should establish some alignment relation “f” between the systems. Generally, as the result, the aligned system A should do "what the system H wants". Two things stand out in this basic framing: Alignment is a relation,
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- ALIGNMENT - by vincent huang - a slice of my mindmindslice.substack.com
- AGI Ruin: A List of Lethalities — LessWronglesswrong.com
- The case against AI alignment — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- What Is The Alignment Problem? — LessWronglesswrong.com
- Another (outer) alignment failure story — AI Alignment Forumalignmentforum.org
- Shah and Yudkowsky on alignment failures — LessWronglesswrong.com
- The Alignment Problem — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- What is AI alignment? - by Adam Jones - BlueDot Impactblog.bluedot.org
- [2605.10310] Positive Alignment: Artificial Intelligence for Human Flourishingarxiv.org