Self-Other Overlap: A Neglected Approach to AI Alignment — LessWrong
Figure 1. Image generated by DALL·E 3 to represent the concept of self-other overlap Many thanks to Bogdan Ionut-Cirstea, Steve Byrnes, Gunnar Zarnac…
x Self-Other Overlap: A Neglected Approach to AI Alignment — LessWrong AI Risk Has Diagram AI Safety Camp Deceptive Alignment Embedded Agency AI Frontpage 247 Self-Other Overlap: A Neglected Approach to AI Alignment by Marc Carauleanu , Mike Vaiana , Kvee , Diogo de Lucena , Cameron Berg , Trent Hodgeson 30th Jul 2024 AI Alignment Forum 14 min read 53 247 Ω 61 Figure 1. Image generated by DALL·E 3 to represent the concept of self-other overlap Many thanks to Bogdan Ionut-Cirstea, Steve Byrnes, Gunnar Zarnacke, Jack Foxabbott and Seong Hah Cho for critical comments and feedback on earlier and o
Explore this link on the map →related reading
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com
- Should We Train Against (CoT) Monitors? — LessWronglesswrong.com
- Deep Deceptiveness — LessWronglesswrong.com
- How likely is deceptive alignment? — AI Alignment Forumalignmentforum.org
- Externalized reasoning oversight: a research direction for language model alignment — AI Alignment Forumalignmentforum.org
- “Alignment Faking” frame is somewhat fake — LessWronglesswrong.com