flâneur — a map of the web's best reading

Realistic Evaluations Will Not Prevent Evaluation Awareness — LessWrong

lesswrong.com · 3,227 words · saved by 1 readers

One Minute Summary I think there's a fundamental limit to behavioral alignment evaluations that gets worse as models improve: Humans control all inpu…

x Realistic Evaluations Will Not Prevent Evaluation Awareness — LessWrong AI Frontpage 38 Realistic Evaluations Will Not Prevent Evaluation Awareness by Adam Karvonen 24th Feb 2026 7 min read 9 38 One Minute Summary I think there's a fundamental limit to behavioral alignment evaluations that gets worse as models improve: Humans control all inputs to a model and can snapshot, replay, or fabricate any scenario at will. An intelligent model could realize this and rationally treat every interaction as a potential evaluation. This means that even perfectly realistic evaluations, including ones samp

Explore this link on the map →

related reading