Catching AIs red-handed — LessWrong
We’ve been spending a lot of time recently thinking about how to mitigate risks posed by scheming (a.k.a. deceptively aligned) models intentionally s…
x Catching AIs red-handed — LessWrong Best of LessWrong 2024 Deceptive Alignment AI Control Redwood Research AI Frontpage 116 Catching AIs red-handed by ryan_greenblatt , Buck 5th Jan 2024 AI Alignment Forum 21 min read 28 116 Ω 62 We’ve been spending a lot of time recently thinking about how to mitigate risks posed by scheming (a.k.a. deceptively aligned) models intentionally subverting your safety techniques in order to cause an unacceptable outcome. We empirically investigated techniques that are robust to intentional subversion in our recent paper . In this post, we’ll discuss a crucial dy
Explore this link on the map →related reading
- Catching AIs red-handedblog.redwoodresearch.org
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- How will we update about scheming? - by Ryan Greenblattredwoodresearch.substack.com
- Buck Shlegeris on controlling AI that wants to take over – so we can use it anyway | 80,000 Hours80000hours.org
- Research Areas in AI Control (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- How will we update about scheming? — LessWronglesswrong.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org
- Current AIs seem pretty misaligned to meblog.redwoodresearch.org