✳flâneur — a map of the web's best reading
Role confusion: sounding like the cause is indistinguishable from being it. — LessWrong
lesswrong.com · 3,861 words · saved by 1 readers
A replication of Prompt Injection as Role Confusion (2026) and why the mechanistic story of prompt injection is harder to pin down than it looks. …
x Role confusion: sounding like the cause is indistinguishable from being it. — LessWrong Interpretability (ML & AI) Jailbreaking (AIs) Role Science AI Frontpage 14 Role confusion: sounding like the cause is indistinguishable from being it. by Owain Mogford 29th Jun 2026 11 min read 9 14 A replication of Prompt Injection as Role Confusion (2026) and why the mechanistic story of prompt injection is harder to pin down than it looks. Epistemic status: I reproduced the direction of the paper's main results on a single consumer GPU (it was faithful in direction but not like for like in magnitude, s
Explore this link on the map →related reading
- [2603.12277] Prompt Injection as Role Confusionarxiv.org
- A Mechanistic Explanation of Prompt Injection (and why you should study roles) — LessWronglesswrong.com
- The Persona Selection Model: Why AI Assistants might Behave like Humansalignment.anthropic.com
- The Waluigi Effect (mega-post) — LessWronglesswrong.com
- How confessions can keep language models honest | OpenAIopenai.com
- llm assistant personas seem increasingly incoherent (some subjective observations) — LessWronglesswrong.com
- Thought Branches: Interpreting LLM Reasoning Requires Resamplingarxiv.org
- confessions_paper.pdfcdn.openai.com
- Role-playing vs Self-modelling — LessWronglesswrong.com
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitorsarxiv.org
- Paper: Prompt Optimization Makes Misalignment Legible — LessWronglesswrong.com
- Discovering Language Model Behaviors with Model-Written Evaluations — LessWronglesswrong.com