Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't scheming
One strategy for mitigating risk from schemers (that is, egregiously misaligned models that intentionally try to subvert your safety measures) is behavioral red-teaming (BRT).
Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't scheming Buck Shlegeris Oct 10, 2024 6 Share One strategy for mitigating risk from schemers (that is, egregiously misaligned models that intentionally try to subvert your safety measures) is behavioral red-teaming (BRT). The basic version of this strategy is something like: Before you deploy your model, but after you train it, you search really hard for inputs on which the model takes actions that are very bad. Then you look at the scariest model actions you found, and if these contain examples that are str
Explore this link on the map →related reading
- Reducing risk from scheming by studying trained-in scheming behavior — LessWronglesswrong.com
- 2312.06942arxiv.org
- Claude Mythos Preview System Cardwww-cdn.anthropic.com
- The case for ensuring that powerful AIs are controlled — LessWronglesswrong.com
- Catching AIs red-handed — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- [2312.06942] AI Control: Improving Safety Despite Intentional Subversionarxiv.org
- AI Control: Improving Safety Despite Intentional Subversion — LessWronglesswrong.com
- Catching AIs red-handedblog.redwoodresearch.org
- How will we update about scheming? - by Ryan Greenblattredwoodresearch.substack.com
- How will we update about scheming? — LessWronglesswrong.com
- Three Sketches of ASL-4 Safety Case Componentsalignment.anthropic.com