A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWrong
lesswrong.com · 12,150 words · saved by 6 readers
Summary * We've been building a theory of how prompt injections work under the hood. * We show it comes down to how LLMs perceive roles (the humble…
x A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWrong Jailbreaking (AIs) Interpretability (ML & AI) Role Science AI Curated 2026 Top Fifty: 14 % 371 A Mechanistic Explanation of Prompt Injection (and why you should study roles) by Charles Ye , Jasmine C. 22nd Jun 2026 19 min read 56 371 Summary We've been building a theory of how prompt injections work under the hood. We show it comes down to how LLMs perceive roles (the humble chat template tags). We use this theory to create new attacks, explain some weird mech interp results, and predict when attacks w
saved by
related reading
- Prompt Injection as Role Confusionrole-confusion.github.io
- [2603.12277] Prompt Injection as Role Confusionarxiv.org
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- Role confusion: sounding like the cause is indistinguishable from being it. — LessWronglesswrong.com
- Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWronglesswrong.com
- Introspection Adapters: Training LLMs to Report Their Learned Behaviorsarxiv.org
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- CaMeL offers a promising new direction for mitigating prompt injection attackssimonwillison.net
- Role-playing vs Self-modelling — LessWronglesswrong.com
- [2607.14111] Introspection Fine-Tuning (IFT): Training Small LLMs to Introspectarxiv.org