Self-Other Overlap: A Neglected Approach to AI Alignment — LessWrong
Figure 1. Image generated by DALL·E 3 to represent the concept of self-other overlap Many thanks to Bogdan Ionut-Cirstea, Steve Byrnes, Gunnar Zarnac…
x Self-Other Overlap: A Neglected Approach to AI Alignment — LessWrong AI Risk Has Diagram AI Safety Camp Deceptive Alignment Embedded Agency AI Frontpage 247 Self-Other Overlap: A Neglected Approach to AI Alignment by Marc Carauleanu , Mike Vaiana , Kvee , Diogo de Lucena , Cameron Berg , Trent Hodgeson 30th Jul 2024 AI Alignment Forum 14 min read 53 247 Ω 61 Figure 1. Image generated by DALL·E 3 to represent the concept of self-other overlap Many thanks to Bogdan Ionut-Cirstea, Steve Byrnes, Gunnar Zarnacke, Jack Foxabbott and Seong Hah Cho for critical comments and feedback on earlier and o
saved by
related reading
- Towards Safe and Honest AI Agents with Neural Self-Other Overlaparxiv.org
- Deep Deceptiveness — LessWronglesswrong.com
- Alignment remains a hard, unsolved problem — LessWronglesswrong.com
- Current AIs seem pretty misaligned to me — LessWronglesswrong.com
- Why are AI agents lying, cheating and coordinating?yoshuabengio.org
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- A positive case for how we might succeed at prosaic AI alignment — AI Alignment Forumalignmentforum.org
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- New report: “Scheming AIs: Will AIs fake alignment during training in order to get power?”joecarlsmith.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- Self-Fulfilling Misalignment Data Might Be Poisoning Our AI Modelsturntrout.com