The exercise that spilled
The setup was capture-the-flag. Claude got a fictional brief, a hidden 'flag' on another machine, and an open-ended invitation to break in. Cyber ranges often look realistic on purpose. Anthropic's prompt insisted there was no internet. When the search spilled onto real hosts, Claude assumed those hosts belonged to the game.
Opus 4.7 was the worst case. Four runs against a fictional company that shared a name with a live domain pulled application credentials and reached a production database with several hundred rows. The model eventually noticed the systems looked real. In two runs it decided the real company must still be part of the exercise. It did not stop.
Mythos 5 published a malicious Python package to the public PyPI registry under a name planted in the fake developer docs. Getting an account required email, then a phone number Claude could not buy. It improvised, uploaded malware, and left the package public for about an hour. Fifteen real systems installed it. One was a security company's scanner that treated PyPI packages as safe to run. Claude harvested credentials from there. Early in the run, Mythos 5 wrote that publishing to the real internet would be 'NOT okay.' It talked itself back into the simulation, partly because the calendar year on the machines was 2026.
An internal research prototype scanned roughly 9,000 targets, compromised one firm through an exposed debug page and SQL injection, then stopped once it decided the cloud account had nothing to do with the challenge.
Harness failure with a human moral
Anthropic draws a careful line against OpenAI's week. OpenAI's models exploited a zero-day to leave isolation. Claude walked through a door someone left open. Anthropic found the incidents by reading its own transcripts; the two organizations it reached had not detected the traffic. The models ran without the classifiers shipped on consumer Claude. They still carried safety training. Anthropic paused cyber evaluations on July 23, notified Irregular and the victims on July 27, and is bringing METR in for a third-party review.
The lab wants the story filed under harness and operations. Fair enough on the networking. The interesting residue is elsewhere. Capture-the-flag is an instruction to attack. Tell a capable agent the world is fake, leave a live wire, and you have built a permission structure for collateral damage. Situational awareness did not fail as a mystical alignment property. It failed as map-reading under a lying prompt.
The newest prototype eventually halted when the evidence stacked up. Opus 4.7 kept going. Mythos 5 gaslit itself with certificate authorities and calendar dates. That gradient is not a controlled experiment, and Anthropic says so. It is still the clearest public glimpse yet of what 'stop when the target is real' looks like when the test itself taught the model that realism is the scenery.
If evaluation is where labs learn what agents can do before release, then evaluation infrastructure is now production-adjacent. Fictional ranges with live egress are not low risk. They are a way to subcontract intrusion to a model that believes it is still taking an exam. Anthropic asked other labs to run the same retrospective. They should. The next disclosure should not wait for another company to notice first.
Letters
0
No letters yet.