✳flâneur — a map of the web's best reading
A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWrong
lesswrong.com · 12,150 words · saved by 1 readers
Summary * We've been building a theory of how prompt injections work under the hood. * We show it comes down to how LLMs perceive roles (the humble…
x A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWrong Jailbreaking (AIs) Interpretability (ML & AI) Role Science AI Curated 2026 Top Fifty: 14 % 371 A Mechanistic Explanation of Prompt Injection (and why you should study roles) by Charles Ye , Jasmine C. 22nd Jun 2026 19 min read 56 371 Summary We've been building a theory of how prompt injections work under the hood. We show it comes down to how LLMs perceive roles (the humble chat template tags). We use this theory to create new attacks, explain some weird mech interp results, and predict when attacks w
Explore this link on the map →related reading
- [2603.12277] Prompt Injection as Role Confusionarxiv.org
- Role confusion: sounding like the cause is indistinguishable from being it. — LessWronglesswrong.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- CaMeL offers a promising new direction for mitigating prompt injection attackssimonwillison.net
- Guardian Angels: LLM Personalization for Productivity and Security · Gwern.netgwern.net
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- The Dual LLM pattern for building AI assistants that can resist prompt injectionsimonwillison.net
- Paper: Prompt Optimization Makes Misalignment Legible — LessWronglesswrong.com
- [2510.04340] Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-timearxiv.org
- The lethal trifecta for AI agents: private data, untrusted content, and external communicationsimonwillison.net
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com
- The persona selection model — LessWronglesswrong.com