Towards Safe and Honest AI Agents with Neural Self-Other Overlap
As AI systems increasingly make critical decisions, deceptive AI poses a significant challenge to trust and safety. We present Self-Other Overlap (SOO) fine-tuning, a promising approach in AI Safety that could substantially improve our ability to build honest artificial intelligence. Inspired by cognitive neuroscience research on empathy, SOO aims to align how AI models represent themselves and others. Our experiments on LLMs with 7B, 27B and 78B parameters demonstrate SOO’s efficacy: deceptive responses of Mistral-7B-Instruct-v0.2 dropped from 73.6% to 17.2% with no observed reduction in general task performance, while in Gemma-2-27b-it and CalmeRys-78B-Orpo-v0.1 deceptive responses were reduced from 100% to 9.3% and 2.7%, respectively, with a small impact on capabilities. In reinforcement learning scenarios, SOO-trained agents showed significantly reduced deceptive behavior. SOO’s focus on contrastive self and other-referencing observations offers strong potential for generalization
Michael Vaiana Affiliation: AE Studio Email: mike@ae.studio Judd Rosenblatt Affiliation: AE Studio Email: judd@ae.studio Cameron Berg Affiliation: AE Studio Email: cameron@ae.studio Diogo Schwerz de Lucena Affiliation: AE Studio Email: diogo@ae.studio Abstract As AI systems increasingly make critical decisions, deceptive AI poses a significant challenge to trust and safety. We present Self-Other Overlap (SOO) fine-tuning, a promising approach in AI Safety that could substantially improve our ability to build honest artificial intelligence. Inspired by cognitive neuroscience…
saved by
related reading
- Self-Other Overlap: A Neglected Approach to AI Alignment — LessWronglesswrong.com
- [2503.03750] The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systemsarxiv.org
- SPAR Research Librarylibrary.sparai.org
- Deep Deceptiveness — LessWronglesswrong.com
- [2401.05566] Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Trainingarxiv.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- How confessions can keep language models honest | OpenAIopenai.com
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org
- [2602.15515] The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probesarxiv.org
- 39 - Evan Hubinger on Model Organisms of Misalignment | AXRP - the AI X-risk Research Podcastaxrp.net
- On Anthropic's Sleeper Agents Paperthezvi.substack.com
- Alignment Faking Mitigationsalignment.anthropic.com