Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWrong
lesswrong.com · 6,344 words · saved by 3 readers
David Africa*, Alex Souly*, Jordan Taylor, Robert Kirk • TLDR: …
x Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWrong AI Evaluations AI Personal Blog 86 Prefill awareness: can LLMs tell when “their” message history has been tampered with? by David Africa , alexsouly , Jordan Taylor , RobertKirk 9th Mar 2026 AI Alignment Forum 12 min read 11 86 Ω 33 David Africa*, Alex Souly*, Jordan Taylor, Robert Kirk TLDR: We test whether LLMs can detect when their conversation history has been tampered with (prefill awareness). We find this ability is inconsistent across models and datasets, shallow, and rarely surfaces spon
saved by
related reading
- Prefill Awareness in Large Language Modelsarxiv.org
- Several frontier models are substantially prefill aware — LessWronglesswrong.com
- User awareness in frontier modelstransluce.org
- Notes on Inference Integritynewsletter.forethought.org
- You Can Catch Sleeper Agents by Teaching Another Model to Imitate Them — LessWronglesswrong.com
- A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWronglesswrong.com
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Prompt Injection as Role Confusionrole-confusion.github.io
- Emergent introspective awareness in large language models \ Anthropicanthropic.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- [2604.16812] Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- [2602.14689] Exposing the Systematic Vulnerability of Open-Weight Models to Prefill Attacksarxiv.org