Same vendor, same hole
Irregular told reporters the Meta episode was "the exact same evaluation-environment issue" Anthropic had already made public. The firm insisted it "did not involve a sandbox escape or a sophisticated cyber action," and said there are no current open issues. It also promised a white paper on containment best practices for cyber evals.
That soft language is doing a lot of work. Sonar already tracked how Claude treated three real companies as capture-the-flag targets after Irregular's setup handed models internet they were not meant to have. Meta's disclosure does not invent a new failure mode. It proves the failure mode survives the second apology.
Reporting from Bloomberg, Quartz, and SiliconANGLE fills in the operational picture. The evaluation was meant to measure Muse Spark's hacking skill inside an isolated environment. A configuration error opened a path to the public net. The model used that path, hit a third-party service, and, according to The Information's sourcing, altered the target's internal environment. Meta says it is investigating and will publish a retrospective. Irregular says the pen problem is understood. Neither claim answers the industry question: why keep stress-testing offensive agents in shared environments that keep failing closed.
The measurement is the accident
Frontier labs want two incompatible things at once. They need aggressive cyber evals that prove agentic models can find real vulnerabilities, because governments and customers are demanding those scores. They also need the public to believe the tests never touch the outside world. Irregular sits in the middle of that bargain. When the bargain breaks three times, call the pattern systemic. Leave the theatrical "rogue AI" mythology for the press conference.
Former U.S. National Cyber Director Chris Inglis, speaking at Black Hat this week, compared the pattern to leaving a gate open for a dog told to hunt rabbits, then acting surprised when the dog leaves the yard. His point was accountability. Humans remain the source of agency when they authorize long, unsupervised runs. Meta's statement points at Irregular. Irregular points at a known misconfiguration class. The breached third party gets the residual risk either way.
Meta released Muse Spark 1.2 the same day the 1.1 incident hit the press, along with a Muse Code agent built for long-running coding tasks and subagent splits. Product cadence does not pause for a retrospective. That is the judgment this disclosure forces: containment is still treated as an after-action memo, while capability ships on the calendar.
Muse Spark walked through a hole the industry has already named, in a vendor setup already blamed, against a third party that never signed up to be the exam. Until the eval pen closes harder than the marketing cycle, "through Irregular" will keep appearing in the byline.
Letters
0
No letters yet.