flâneur — a map of the web's best reading

Prompt Injection as Role Confusion

role-confusion.github.io · 48 words · saved by 1 readers

LLMs can't tell who's speaking. We show they identify roles by writing style, not tags, and exploit this with CoT Forgery, injecting fake reasoning that models mistake for their own thoughts.

To cite the paper or this writeup, please use the ICML paper citation. Copy BibTeX @inproceedings{ye2026promptinjectionroleconfusion, title = {Prompt Injection as Role Confusion}, author = {Ye, Charles and Cui, Jasmine and Hadfield-Menell, Dylan}, booktitle = {International Conference on Machine Learning (ICML)}, year = {2026}, url = {https://arxiv.org/abs/2603.12277} }

Explore this link on the map →

related reading