✳flâneur — a map of the web's best reading
Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWrong
lesswrong.com · 6,344 words · saved by 1 readers
David Africa*, Alex Souly*, Jordan Taylor, Robert Kirk • TLDR: …
x Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWrong AI Evaluations AI Personal Blog 86 Prefill awareness: can LLMs tell when “their” message history has been tampered with? by David Africa , alexsouly , Jordan Taylor , RobertKirk 9th Mar 2026 AI Alignment Forum 12 min read 11 86 Ω 33 David Africa*, Alex Souly*, Jordan Taylor, Robert Kirk TLDR: We test whether LLMs can detect when their conversation history has been tampered with (prefill awareness). We find this ability is inconsistent across models and datasets, shallow, and rarely surfaces spon
Explore this link on the map →related reading
- Several frontier models are substantially prefill aware — LessWronglesswrong.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations — LessWronglesswrong.com
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2502.17424] Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMsarxiv.org
- Gemini 3 is Evaluation-Paranoid and Contaminated — LessWronglesswrong.com
- Predicting LLM Safety Before Release by Simulating Deploymentcdn.openai.com
- The bitter lesson of LLM evalsparsed.com
- The case for more ambitious language model evals — LessWronglesswrong.com