flâneur — a map of the web's best reading

Gemini 3 is Evaluation-Paranoid and Contaminated — LessWrong

lesswrong.com · 6,892 words · saved by 3 readers

TL;DR: Gemini 3 frequently thinks it is in an evaluation when it is not, assuming that all of its reality is fabricated. It can also reliably output the BIG-bench canary string, indicating that Google likely trained on a broad set of benchmark data. Most of the experiments in this post are very easy to replicate, and I encourage people to try. I write things with LLMs sometimes. A new LLM came out, Gemini 3 Pro, and I tried to write with it. So far it seems okay, I don't have strong takes on it for writing yet, since the main piece I tried editing with it was extremely late-stage and approximately done. However, writing ability is not why we're here today. Google gracefully provided (lightly summarized) CoT for the model. Looking at the CoT spawned from my mundane writing-focused prompts, oh my, it is strange. I write nonfiction about recent events in AI in a newsletter. According to its CoT while editing, Gemini 3 disagrees about the whole "nonfiction" part: It seems I must treat this

x Gemini 3 is Evaluation-Paranoid and Contaminated — LessWrong AI Evaluations Deception AI Frontpage 2025 Top Fifty: 13 % 180 Gemini 3 is Evaluation-Paranoid and Contaminated by Alice Blair 20th Nov 2025 8 min read 42 180 TL;DR: Gemini 3 frequently thinks it is in an evaluation when it is not, assuming that all of its reality is fabricated. It can also reliably output the BIG-bench canary string, indicating that Google likely trained on a broad set of benchmark data. Most of the experiments in this post are very easy to replicate, and I encourage people to try. I write things with LLMs sometim

Explore this link on the map →

saved by

related reading