Role confusion: sounding like the cause is indistinguishable from being it. — LessWrong
lesswrong.com · 3,861 words · saved by 1 readers
A replication of Prompt Injection as Role Confusion (2026) and why the mechanistic story of prompt injection is harder to pin down than it looks. …
x Role confusion: sounding like the cause is indistinguishable from being it. — LessWrong Interpretability (ML & AI) Jailbreaking (AIs) Role Science AI Frontpage 14 Role confusion: sounding like the cause is indistinguishable from being it. by Owain Mogford 29th Jun 2026 11 min read 9 14 A replication of Prompt Injection as Role Confusion (2026) and why the mechanistic story of prompt injection is harder to pin down than it looks. Epistemic status: I reproduced the direction of the paper's main results on a single consumer GPU (it was faithful in direction but not like for like in magnitude, s
related reading
- A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWronglesswrong.com
- Prompt Injection as Role Confusionrole-confusion.github.io
- [2603.12277] Prompt Injection as Role Confusionarxiv.org
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- confessions_paper.pdfcdn.openai.com
- Why We Are Excited About Confessionsalignment.openai.com
- Prefill awareness: can LLMs tell when “their” message history has been tampered with? — LessWronglesswrong.com
- Thought Branches: Interpreting LLM Reasoning Requires Resamplingarxiv.org
- Agentic Misalignment in Summer 2026alignment.anthropic.com
- [2606.26071] Model Forensics: Investigating Whether Concerning Behavior Reflects Misalignmentarxiv.org