flâneur — a map of the web's best reading

Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't scheming

blog.redwoodresearch.org · 3,950 words · saved by 1 readers

One strategy for mitigating risk from schemers (that is, egregiously misaligned models that intentionally try to subvert your safety measures) is behavioral red-teaming (BRT).

Behavioral red-teaming is unlikely to produce clear, strong evidence that models aren't scheming Buck Shlegeris Oct 10, 2024 6 Share One strategy for mitigating risk from schemers (that is, egregiously misaligned models that intentionally try to subvert your safety measures) is behavioral red-teaming (BRT). The basic version of this strategy is something like: Before you deploy your model, but after you train it, you search really hard for inputs on which the model takes actions that are very bad. Then you look at the scariest model actions you found, and if these contain examples that are str

Explore this link on the map →

related reading