flâneur — a map of the web's best reading

Models May Behave Worse When Eval Aware — LessWrong

lesswrong.com · 10,218 words · saved by 1 readers

This is the first in a series of research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent are…

x Models May Behave Worse When Eval Aware — LessWrong AI Frontpage 90 Models May Behave Worse When Eval Aware by Senthooran Rajamanoharan , Neel Nanda 11th Jun 2026 AI Alignment Forum 16 min read 8 90 Ω 37 This is the first in a series of research updates from the Google DeepMind Language Model Interpretability team, in interpretability and adjacent areas. TL;DR It's often assumed that models will act more aligned when they can tell they're being evaluated. But we find that Gemini can take “undesired” actions in behavioural evals even when it explicitly reasons that the environments are contri

Explore this link on the map →

related reading