Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't scheming
One strategy for mitigating risk from schemers (that is, egregiously misaligned models that intentionally try to subvert your safety measures) is behavioral red-teaming (BRT).
Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't scheming Buck Shlegeris Oct 10, 2024 6 Share One strategy for mitigating risk from schemers (that is, egregiously misaligned models that intentionally try to subvert your safety measures) is behavioral red-teaming (BRT). The basic version of this strategy is something like: Before you deploy your model, but after you train it, you search really hard for inputs on which the model takes actions that are very bad. Then you look at the scariest model actions you found, and if these contain examples that are str
related reading
- Reducing risk from scheming by studying trained-in scheming behavior — LessWronglesswrong.com
- 2312.06942arxiv.org
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- How will we update about scheming?blog.redwoodresearch.org
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Catching AIs red-handed — LessWronglesswrong.com
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- We need 3rd party Training-Run Assessments — LessWronglesswrong.com
- Many arguments for AI x-risk are wrong — AI Alignment Forumalignmentforum.org
- Catching AIs red-handedblog.redwoodresearch.org