Gemini 3 is Evaluation-Paranoid and Contaminated — LessWrong
TL;DR: Gemini 3 frequently thinks it is in an evaluation when it is not, assuming that all of its reality is fabricated. It can also reliably output the BIG-bench canary string, indicating that Google likely trained on a broad set of benchmark data. Most of the experiments in this post are very easy to replicate, and I encourage people to try. I write things with LLMs sometimes. A new LLM came out, Gemini 3 Pro, and I tried to write with it. So far it seems okay, I don't have strong takes on it for writing yet, since the main piece I tried editing with it was extremely late-stage and approximately done. However, writing ability is not why we're here today. Google gracefully provided (lightly summarized) CoT for the model. Looking at the CoT spawned from my mundane writing-focused prompts, oh my, it is strange. I write nonfiction about recent events in AI in a newsletter. According to its CoT while editing, Gemini 3 disagrees about the whole "nonfiction" part: It seems I must treat this
x Gemini 3 is Evaluation-Paranoid and Contaminated — LessWrong AI Evaluations Deception AI Frontpage 2025 Top Fifty: 13 % 180 Gemini 3 is Evaluation-Paranoid and Contaminated by Alice Blair 20th Nov 2025 8 min read 42 180 TL;DR: Gemini 3 frequently thinks it is in an evaluation when it is not, assuming that all of its reality is fabricated. It can also reliably output the BIG-bench canary string, indicating that Google likely trained on a broad set of benchmark data. Most of the experiments in this post are very easy to replicate, and I encourage people to try. I write things with LLMs sometim
Explore this link on the map →saved by
related reading
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- gpt-4.pdfcdn.openai.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWronglesswrong.com
- Models May Behave Worse When Eval Aware — LessWronglesswrong.com
- The bitter lesson of LLM evalsparsed.com
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations – Apollo Researchapolloresearch.ai
- We need a better way to evaluate emergent misalignment — LessWronglesswrong.com
- Predicting LLM Safety Before Release by Simulating Deploymentcdn.openai.com
- LLM Evaluation doesn't need to be complicatedphilschmid.de