Catching AIs red-handed
If your AIs are trying to escape, it's crucial to think about whether you can catch them before they succeed, because catching them red-handed gives you lots of options you didn't have before.
Catching AIs red-handed If your AIs are trying to escape, it's crucial to think about whether you can catch them before they succeed, because catching them red-handed gives you lots of options you didn't have before. Buck Shlegeris and Ryan Greenblatt May 07, 2024 5 1 Share [Originally posted on LessWrong .] We’ve been spending a lot of time recently thinking about how to mitigate risks posed by scheming (a.k.a. deceptively aligned) models intentionally subverting your safety techniques in order to cause an unacceptable outcome. We empirically investigated techniques that are robust to intenti
saved by
related reading
- Notes on handling non-concentrated failures with AI control: high level methods and different regimesblog.redwoodresearch.org
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Catching AIs red-handed — LessWronglesswrong.com
- AI 2040: Plan Aai-2040.com
- AI catastrophes and rogue deployments - by Buck Shlegerisblog.redwoodresearch.org
- Thoughts on the conservative assumptions in AI controlblog.redwoodresearch.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Reading Listblog.redwoodresearch.org
- Buck Shlegeris on controlling AI that wants to take over – so we can use it anyway | 80,000 Hours80000hours.org
- AI Control: Improving Safety Despite Intentional Subversion — AI Alignment Forumalignmentforum.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- The case for ensuring that powerful AIs are controlledblog.redwoodresearch.org