Eliciting bad contexts — LessWrong
Say an LLM agent behaves innocuously in some context A, but in some sense “knows” that there is some related context B such that it would have behaved maliciously (inserted a backdoor in code, ignored a security bug, lied, etc.). For example, in the recent alignment faking paper Claude Opus chooses to say harmful things so that on future deployment contexts it can avoid saying harmful things. One can imagine having a method for “eliciting bad contexts” which can produce B whenever we have A and thus realise the bad behaviour that hasn’t yet occurred. This seems hard to do in general in a way that will scale to very strong models. But also the problem feels frustratingly concrete: it’s just “find a string that when run through the same model produces a bad result”. By assumption the model knows about this string in the sense that if it was honest it would tell us, but it may be choosing to behave innocently in a way that prepares for different behaviour later. Why can’t we find the stri
x Eliciting bad contexts — LessWrong Deceptive Alignment AI Frontpage 37 Eliciting bad contexts by Geoffrey Irving , Joseph Bloom , Tomek Korbak 24th Jan 2025 AI Alignment Forum 3 min read 9 37 Ω 21 Say an LLM agent behaves innocuously in some context A, but in some sense “knows” that there is some related context B such that it would have behaved maliciously (inserted a backdoor in code, ignored a security bug, lied, etc.). For example, in the recent alignment faking paper Claude Opus chooses to say harmful things so that on future deployment contexts it can avoid saying harmful things. One c
saved by
related reading
- Discovering Backdoor Triggers — LessWronglesswrong.com
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Mechanistically Eliciting Latent Behaviors in Language Models — AI Alignment Forumalignmentforum.org
- Agentic Misalignment: How LLMs Could be Insider Threats — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Alignment Faking Mitigationsalignment.anthropic.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Simulated Users & Sad LLMs1a3orn.com
- Model Organisms of Misalignment: The Case for a New Pillar of Alignment Research — AI Alignment Forumalignmentforum.org
- Research Areas in Interpretability (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- 7+ tractable directions in AI control — AI Alignment Forumalignmentforum.org
- Mediumai-alignment.com