AI Control: Improving Safety Despite Intentional Subversion — LessWrong
We’ve released a paper, AI Control: Improving Safety Despite Intentional Subversion. This paper explores techniques that prevent AI catastrophes even…
x AI Control: Improving Safety Despite Intentional Subversion — LessWrong Best of LessWrong 2023 AI Control Deceptive Alignment Redwood Research AI Alignment Intro Materials AI Frontpage 240 AI Control: Improving Safety Despite Intentional Subversion by Buck , Fabien Roger , ryan_greenblatt , Kshitij Sachan 13th Dec 2023 AI Alignment Forum 12 min read 26 240 Ω 102 We’ve released a paper, AI Control: Improving Safety Despite Intentional Subversion . This paper explores techniques that prevent AI catastrophes even if AI instances are colluding to subvert the safety techniques. In this post: We s
Explore this link on the map →saved by
related reading
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- 2312.06942arxiv.org
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionar5iv.labs.arxiv.org
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org
- AI catastrophes and rogue deployments - by Buck Shlegerisblog.redwoodresearch.org