This is a masterful demonstration of the unsecurable nature of LLMs, and how prompt injection by dedicated humans who know how to write will always win.
Prompt Injection as Role Confusion
LLMs can't tell who's speaking. We show they identify roles by writing style, not tags, and exploit this with CoT Forgery, injecting fake reasoning that models mistake for their own thoughts.role-confusion.github.io
