LessWrong
Some things are fundamentally "out to get you," seeking your time, money and attention. Slack allows margin of error. You can relax. You can explore and pursue opportunities. You can plan for the long term. You can stick to principles and do the right thing. Back in 2020, @Raemon gave me some extremely good advice. @johnswentworth had left some comments on a post of mine that I found extremely frustrating and counterproductive. At the time I had no idea about his body of work, so he was just some annoying guy. Ray, who did know who John was and thought he was doing important work, told me: You can't save the world without working with people at least as annoying as John. Which didn't mean I had to heal the rift with John in particular, but if I was going to make that a policy then I would need to give up on my goal of having real impact. John and I did a video call, and it went well. He pointed out a major flaw in my post, I impressed him by immediately updating once he pointed it out
x Gemini 3 is Evaluation-Paranoid and Contaminated — LessWrong AI Evaluations Deception AI Frontpage 2025 Top Fifty: 13 % 181 Gemini 3 is Evaluation-Paranoid and Contaminated by Alice Blair 20th Nov 2025 8 min read 42 181 TL;DR: Gemini 3 frequently thinks it is in an evaluation when it is not, assuming that all of its reality is fabricated. It can also reliably output the BIG-bench canary string, indicating that Google likely trained on a broad set of benchmark data. Most of the experiments in this post are very easy to replicate, and I encourage people to try. I write things with LLMs sometim
Explore this link on the map →related reading
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com
- gpt-4.pdfcdn.openai.com
- trees are harlequins, words are harlequins - the voidnostalgebraist.tumblr.com
- Modifying LLM Beliefs with Synthetic Document Finetuningalignment.anthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Models May Behave Worse When Eval Aware — LessWronglesswrong.com
- The bitter lesson of LLM evalsparsed.com
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWronglesswrong.com
- 2025: The year in LLMssimonwillison.net
- Why Do Naive SFT Filters For Safety Properties Fail? — AI Alignment Forumalignmentforum.org
- Predicting LLM Safety Before Release by Simulating Deploymentcdn.openai.com