Models May Behave Worse When Eval Aware — LessWrong
This is the first in a series of research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent are…
x Models May Behave Worse When Eval Aware — LessWrong AI Frontpage 90 Models May Behave Worse When Eval Aware by Senthooran Rajamanoharan , Neel Nanda 11th Jun 2026 AI Alignment Forum 16 min read 8 90 Ω 37 This is the first in a series of research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. TL;DR It's often assumed that models will act more aligned when they can tell they're being evaluated. But we find that Gemini can take “undesired” actions in behavioural evals even when it explicitly reasons that the environments are contri
Explore this link on the map →related reading
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWronglesswrong.com
- Teaching Claude why \ Anthropicanthropic.com
- Natural-emergent-misalignment-from-reward-hacking-paper.pdfassets.anthropic.com
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com
- Metagaming matters for training, evaluation, and oversightalignment.openai.com
- Realistic Evaluations Will Not Prevent Evaluation Awareness — LessWronglesswrong.com
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations – Apollo Researchapolloresearch.ai
- A Toy Environment For Exploring Reasoning About Reward — LessWronglesswrong.com
- How to Design Environments for Understanding Model Motives — LessWronglesswrong.com
- Reproducing steering against evaluation awareness in a large open-weight model — LessWronglesswrong.com
- How Can Interpretability Researchers Help AGI Go Well? — AI Alignment Forumalignmentforum.org