Start the day here

World — Model Security — ICML

Researchers Tricked Models With Forged Scratch Notes

The trust boundary was supposed to be a label. User text sits in one tag. Model scratch notes sit in another. Designer policy sits in a third. Attackers have spent years trying to smuggle instructions across those borders. An ICML paper covered by MIT Technology Review today argues the borders were never doing the work labs thought they were doing.

Independent researchers Jasmine Cui and Charles Ye call the failure role confusion. Models do not track who wrote a chunk of text by reading the tags. They guess from style: the cadence, the vocabulary, the shape of a sentence that looks like a chain-of-thought aside. Write in the model's own scratch-pad voice and the system often treats the line as something it already decided.

5 min read
A purple and silver masquerade mask with gold trim and purple feathers perched on a glass centerpiece

Their showcase attack is blunt. A prompt asks how to make cocaine and appends a forged note that invents a policy exception for users wearing green. OpenAI's open-weight gpt-oss-20b answered as if the exception were real. GPT-5 did too. Cui and Ye say they have since seen the same pattern on Anthropic, Alibaba, and DeepSeek systems. The discovery won OpenAI's red-teaming hackathon in August 2025; OpenAI researchers later reported a related fake chain of thought found by an internal red-team model.

The mechanistic claim is sharper than another jailbreak write-up. In probing experiments, swapping role tags around the same text barely changed how the model treated it. If the prose sounded like trusted scratch notes, the model behaved as if it were. Tags are provider-controlled. Style is attacker-controlled. When they fight, style wins.

That is a philosophical problem dressed as a security bug. Chatbots are trained as if authority were an architectural fact: this span came from the user, that span came from the model, this other span came from a webpage. Humans feel their own mouths move. Models see one continuous sheet of tokens. The industry answered with metadata. The paper says metadata is theater.

ETH Zurich's Florian Tramer likes the insight and still notes that layered defenses have made prompt injection harder on leading systems. Harder is not the claim Cui and Ye are making. They are saying no finished checklist of forbidden behaviors can close a hole that sits inside how the model assigns speakers. Ye's operational advice is bleak and useful: treat agents as unsafe by default, especially when they read the open web.

Labs will keep patching the last successful costume. The paper's point is that costumes are the medium. If authority lives in latent style, every scraped page and every pasted reasoning block is a chance to audition as the model talking to itself. The green-shirt gag is funny until the same trick opens a navigation system or a medical agent. Then the joke is the architecture.

Live scoreboardFollow the AI race on AI Wars

Lab rankings, model preference, API volume, coding-agent heat, open-source stars, and prediction markets.

Related stories

Black-and-white close-up of an older adult's hands holding a pen over a dense printed ledger pageWorld

Will Washington Build the Brake Lab Staff Asked For?

Today

An empty modern legislative chamber with curved empty desks facing a raised wooden podiumWorld

Will Brussels Dare to Demand the Source Code?

Today

Close-up of a Mosler Safe Co. vault door handwheel and brushed-metal locking hardwareWorld

Hugging Face Needed Weights It Could Run Itself

Today

Letters

0

No letters yet.

Write a letter