Catching AIs red-handed — LessWrong
We’ve been spending a lot of time recently thinking about how to mitigate risks posed by scheming (a.k.a. deceptively aligned) models intentionally s…
x Catching AIs red-handed — LessWrong Best of LessWrong 2024 Deceptive Alignment AI Control Redwood Research AI Frontpage 116 Catching AIs red-handed by ryan_greenblatt , Buck 5th Jan 2024 AI Alignment Forum 21 min read 28 116 Ω 62 We’ve been spending a lot of time recently thinking about how to mitigate risks posed by scheming (a.k.a. deceptively aligned) models intentionally subverting your safety techniques in order to cause an unacceptable outcome. We empirically investigated techniques that are robust to intentional subversion in our recent paper . In this post, we’ll discuss a crucial dy
related reading
- Catching AIs red-handedblog.redwoodresearch.org
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Reading Listblog.redwoodresearch.org
- How will we update about scheming?blog.redwoodresearch.org
- How will we update about scheming? - by Ryan Greenblattredwoodresearch.substack.com
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- Buck Shlegeris on controlling AI that wants to take over – so we can use it anyway | 80,000 Hours80000hours.org
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- Deep Deceptiveness — LessWronglesswrong.com
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org