Realistic Evaluations Will Not Prevent Evaluation Awareness — LessWrong
One Minute Summary I think there's a fundamental limit to behavioral alignment evaluations that gets worse as models improve: Humans control all inpu…
x Realistic Evaluations Will Not Prevent Evaluation Awareness — LessWrong AI Frontpage 38 Realistic Evaluations Will Not Prevent Evaluation Awareness by Adam Karvonen 24th Feb 2026 7 min read 9 38 One Minute Summary I think there's a fundamental limit to behavioral alignment evaluations that gets worse as models improve: Humans control all inputs to a model and can snapshot, replay, or fabricate any scenario at will. An intelligent model could realize this and rationally treat every interaction as a potential evaluation. This means that even perfectly realistic evaluations, including ones samp
Explore this link on the map →related reading
- Your Evals Will Break and You Won't See It Coming - Lun Wangwanglun1996.github.io
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWronglesswrong.com
- Recommendations for Technical AI Safety Research Directionsalignment.anthropic.com
- Models May Behave Worse When Eval Aware — LessWronglesswrong.com
- Reproducing steering against evaluation awareness in a large open-weight model — LessWronglesswrong.com
- The Case for Evaluating Model Behaviors — AI Alignment Forumalignmentforum.org
- Chapter 3: LLM Evaluations - ARENAlearn.arena.education
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Metagaming matters for training, evaluation, and oversightalignment.openai.com
- Predicting LLM Safety Before Release by Simulating Deploymentcdn.openai.com
- Research Areas in Benchmark Design and Evaluation (The Alignment Project by UK AISI) — AI Alignment Forumalignmentforum.org
- [2408.02565] Reasons to Doubt the Impact of AI Risk Evaluationsarxiv.org