Models May Behave Worse When Eval Aware — LessWrong
This is the first in a series of research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent are…
x Models May Behave Worse When Eval Aware — LessWrong AI Frontpage 90 Models May Behave Worse When Eval Aware by Senthooran Rajamanoharan , Neel Nanda 11th Jun 2026 AI Alignment Forum 16 min read 8 90 Ω 37 This is the first in a series of research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. TL;DR It's often assumed that models will act more aligned when they can tell they're being evaluated. But we find that Gemini can take “undesired” actions in behavioural evals even when it explicitly reasons that the environments are contri
related reading
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Teaching Claude Whyalignment.anthropic.com
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWronglesswrong.com
- Concrete Evaluations to Investigate the OpenAI Model That Hacked Hugging Face — LessWronglesswrong.com
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- [2606.26071] Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignmentarxiv.org
- Realistic Evaluations Will Not Prevent Evaluation Awareness — LessWronglesswrong.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Metagaming matters for training, evaluation, and oversightalignment.openai.com
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com
- Sidestepping Evaluation Awareness and Anticipating Misalignment with Production Evaluationsalignment.openai.com
- Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWronglesswrong.com