flâneur — a map of the web's best reading

Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWrong

lesswrong.com · 6,344 words · saved by 1 readers

David Africa*, Alex Souly*, Jordan Taylor, Robert Kirk • TLDR: …

x Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWrong AI Evaluations AI Personal Blog 86 Prefill awareness: can LLMs tell when “their” message history has been tampered with? by David Africa , alexsouly , Jordan Taylor , RobertKirk 9th Mar 2026 AI Alignment Forum 12 min read 11 86 Ω 33 David Africa*, Alex Souly*, Jordan Taylor, Robert Kirk TLDR: We test whether LLMs can detect when their conversation history has been tampered with (prefill awareness). We find this ability is inconsistent across models and datasets, shallow, and rarely surfaces spon

Explore this link on the map →

related reading