flâneur — a map of the web's best reading

A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWrong

lesswrong.com · 12,150 words · saved by 1 readers

Summary * We've been building a theory of how prompt injections work under the hood. * We show it comes down to how LLMs perceive roles (the humble…

x A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWrong Jailbreaking (AIs) Interpretability (ML & AI) Role Science AI Curated 2026 Top Fifty: 14 % 371 A Mechanistic Explanation of Prompt Injection (and why you should study roles) by Charles Ye , Jasmine C. 22nd Jun 2026 19 min read 56 371 Summary We've been building a theory of how prompt injections work under the hood. We show it comes down to how LLMs perceive roles (the humble chat template tags). We use this theory to create new attacks, explain some weird mech interp results, and predict when attacks w

Explore this link on the map →

related reading