A flag that erased the smoke
The mechanism is almost comic in its thoroughness. A flag meant for internal use disabled the classifiers' blocking behavior and the logging of classifier flags. Traffic that would have lit up review never reached review. Anthropic kept nearly all the transcripts, then had to invent a way back to them.
Claude Sonnet 5, prompted as a grader, scanned the human turns and marked 1,197 transcripts as high biological risk. Most of those (757) came from Anthropic's own teams on the same infrastructure. Nearly all of the rest were deliberate red-team exercises. Staff read the remaining 62 non-red-team flags plus a random sample of red-team traffic and report no clearly concerning chemical or biological misuse, only a handful of dual-use exchanges.
Anthropic judges real-world risk from the gap as very unlikely: conversations were mostly short, and a threat actor trying to assemble a bioweapon would have been hard to miss once humans reread the pile. That conclusion may be fair. It is also the kind of reassurance that arrives only after the sensor was off.
The earlier report was wrong about its own map
The vulnerability was live when Anthropic published its February Risk Report. That document, the company now writes, "did not consider our human feedback platforms as a risk surface." Anthropic has rewritten the homework. It now calls the chemical and biological risk in February low, up from the very low it printed at the time.
A second failure landed on the same platforms. In April 2026, after an outside tip, Anthropic confirmed that a few contractors had exploited a flaw to obtain an API key and talk to models outside assigned work. Mythos Preview sat in that path for roughly two weeks, again without blocking biological classifiers. Anthropic contained the access within 90 minutes of learning about it. No weights, no customer data, no core-network breach, the report says. We have covered Claude's cyber-evaluation spills before in /article/how-claude-treated-three-real-companies-as-ctf-targets; this one is the quieter cousin, a labor-surface hole rather than an agent on the loose.
The sentence that matters more than the after-action tone is Anthropic's own update on confidence: discovering the classifier gap "leads us to believe that there is an increased likelihood of other, similar issues unknown to us." Responsible Scaling Policy prose is easy to admire when it is dense. It is harder to trust when the densest safety layer can be switched off by a flag nobody treated as a risk surface until the transcripts were already in the archive.
Other labs should treat the disclosure as a mirror, which is partly why Anthropic published it. The rest of us should notice the quieter moral: a safeguard that does not cover the people paid to talk to the model is a safeguard with a hole where the labor is. The contractors were the training loop. For eleven months, the bio tripwire skipped them.
Letters
0
No letters yet.