Prompt Injection as Role Confusion
LLMs can't tell who's speaking. We show they identify roles by writing style, not tags, and exploit this with CoT Forgery, injecting fake reasoning that models mistake for their own thoughts.
Extended writeup (June 2026) A Theory of Prompt Injection (and why you should study roles) This is a blog-style writeup of the paper. We show prompt injections are driven by a flaw in how LLMs perceive roles. This lets us create new attacks, explain mech interp results, and predict when attacks succeed. We then discuss what roles are and why they matter, and share research ideas for a science of roles. 1. The World to an LLM How does an LLM know the difference between its own thoughts and someone else's words? To see why this is hard, let's look at what the world actually looks like to…
saved by
related reading
- A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWronglesswrong.com
- [2603.12277] Prompt Injection as Role Confusionarxiv.org
- [2603.07267] How to Steal Reasoning Without Reasoning Tracesarxiv.org
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Role confusion: sounding like the cause is indistinguishable from being it. — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- Emergent introspective awareness in large language models \ Anthropicanthropic.com
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- llm-security/README.md at main · greshake/llm-securitygithub.com
- LLM Powered Autonomous Agents | Lil'Loglilianweng.github.io
- Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWronglesswrong.com