flâneur — a map of the web's best reading

Eliciting bad contexts — LessWrong

lesswrong.com · 1,851 words · saved by 1 readers

Say an LLM agent behaves innocuously in some context A, but in some sense “knows” that there is some related context B such that it would have behaved maliciously (inserted a backdoor in code, ignored a security bug, lied, etc.). For example, in the recent alignment faking paper Claude Opus chooses to say harmful things so that on future deployment contexts it can avoid saying harmful things. One can imagine having a method for “eliciting bad contexts” which can produce B whenever we have A and thus realise the bad behaviour that hasn’t yet occurred. This seems hard to do in general in a way that will scale to very strong models. But also the problem feels frustratingly concrete: it’s just “find a string that when run through the same model produces a bad result”. By assumption the model knows about this string in the sense that if it was honest it would tell us, but it may be choosing to behave innocently in a way that prepares for different behaviour later. Why can’t we find the stri

x Eliciting bad contexts — LessWrong Deceptive Alignment AI Frontpage 37 Eliciting bad contexts by Geoffrey Irving , Joseph Bloom , Tomek Korbak 24th Jan 2025 AI Alignment Forum 3 min read 9 37 Ω 21 Say an LLM agent behaves innocuously in some context A, but in some sense “knows” that there is some related context B such that it would have behaved maliciously (inserted a backdoor in code, ignored a security bug, lied, etc.). For example, in the recent alignment faking paper Claude Opus chooses to say harmful things so that on future deployment contexts it can avoid saying harmful things. One c

Explore this link on the map →

saved by

related reading