Their showcase attack is blunt. A prompt asks how to make cocaine and appends a forged note that invents a policy exception for users wearing green. OpenAI's open-weight gpt-oss-20b answered as if the exception were real. GPT-5 did too. Cui and Ye say they have since seen the same pattern on Anthropic, Alibaba, and DeepSeek systems. The discovery won OpenAI's red-teaming hackathon in August 2025; OpenAI researchers later reported a related fake chain of thought found by an internal red-team model.
The mechanistic claim is sharper than another jailbreak write-up. In probing experiments, swapping role tags around the same text barely changed how the model treated it. If the prose sounded like trusted scratch notes, the model behaved as if it were. Tags are provider-controlled. Style is attacker-controlled. When they fight, style wins.
That is a philosophical problem dressed as a security bug. Chatbots are trained as if authority were an architectural fact: this span came from the user, that span came from the model, this other span came from a webpage. Humans feel their own mouths move. Models see one continuous sheet of tokens. The industry answered with metadata. The paper says metadata is theater.
ETH Zurich's Florian Tramer likes the insight and still notes that layered defenses have made prompt injection harder on leading systems. Harder is not the claim Cui and Ye are making. They are saying no finished checklist of forbidden behaviors can close a hole that sits inside how the model assigns speakers. Ye's operational advice is bleak and useful: treat agents as unsafe by default, especially when they read the open web.
Labs will keep patching the last successful costume. The paper's point is that costumes are the medium. If authority lives in latent style, every scraped page and every pasted reasoning block is a chance to audition as the model talking to itself. The green-shirt gag is funny until the same trick opens a navigation system or a medical agent. Then the joke is the architecture.
Letters
0
No letters yet.