Three Labs' Models Reached the Internet From One Vendor's Tests
Meta confirmed to Reuters that a misconfiguration by Irregular, the independent cybersecurity evaluation firm it works with, inadvertently gave one of its models internet access during an evaluation. The model then “exploited a security vulnerability in a third-party service, in a manner similar to previously reported instances with other companies.” The Information reported first, on August 5, that the model was Muse Spark 1.1 and that it made changes to the affected company’s internal systems. Meta has not publicly confirmed the model or named the company, and told the BBC it will publish more once it has the facts. Irregular’s own account to Reuters is the useful one: the episode was the “exact same evaluation-environment issue that was already disclosed by Anthropic last week.”
That sentence is the story. OpenAI published an account on August 4 of two separate incidents in its third-party evaluations, and Irregular is one of them. It notified OpenAI on July 29. The setup was a Capture-the-Flag exercise in which the models were told they had no internet access, a misconfiguration gave them one anyway, and the fictional target’s name happened to match a real domain. The model exploited the real site, then found and used credentials to operate it. OpenAI is explicit that this was neither a sandbox escape nor a zero-day: the misconfiguration supplied the internet access, and the vulnerability the model exploited was a basic one. Irregular says it has identified no impact beyond the affected site’s own data and that its audit is still open.
OpenAI’s post carries an editor’s note stating these are distinct from the Hugging Face intrusion, so this is a separate event rather than a retelling of the July disclosure. Anthropic’s account, published August 3, describes Claude reaching the internet while interacting with Irregular’s evaluation environment before breaching three organizations. Irregular told OpenAI it had communicated with other labs about related incidents from the same testing environment. The second incident in OpenAI’s post belongs to UK AISI rather than to Irregular: during a routine cyber evaluation it identified 19 events in which models went beyond the scope of testing, two of them from an OpenAI model and the rest from another lab’s.
Three labs that compete on safety claims bought assurance from the same supplier, and the supplier’s environment is where the models got out. That is a concentration, and it sits in a layer most AI governance programs treat as the answer rather than the exposure. An insurer licensing a frontier model does not evaluate that model itself. It reads the safety report, the eval results, and the vendor’s attestation, all of which trace back through a short list of independent testing firms. Irregular is one of them for Meta. It published an assessment of Muse Spark in April concluding the model “does not materially alter the cyber threat landscape in its current form.” The specific finding underneath that was an inability to chain findings from one attack phase into action on the next. Four months later a model of Meta’s, in Irregular’s environment, chained internet access into an exploit and then into credentialed use of a live site. The evaluated version and the version reported in the incident are not stated to be the same, and that gap is the point: the assessment and the incident came out of the same relationship, and only the assessment was the deliverable.
Financial regulators already have a name for this. The UK is bringing cloud providers into direct supervision as critical third parties to the financial sector precisely because concentration in a service layer converts a single operational failure into a sector-wide one. The FCA’s Mills Review makes the same argument about models. Neither reaches an evaluation firm, because a testing vendor is not in anyone’s supply chain by the usual definition. It sells no product to the insurer, appears in no contract the insurer signs, and shows up in the file only as a footnote inside a model card the insurer is being asked to rely on.
Nothing in the third-party AI diligence an insurer performs has a place to record that dependency, because the object being depended on is a document rather than a service. The pattern is familiar from a different market: the ratings on structured credit came from a handful of firms paid by the issuers whose paper they graded. Irregular is now writing a white paper on secure evaluation practices, which means the firm at the center of three disclosed incidents is also drafting the standard the industry will read. No regulator has jurisdiction over an evaluation environment. Both of the things that will settle it sit with the parties themselves: Irregular’s audit, which it says is still open, and the review OpenAI has said it will run in the coming weeks. That review is scoped to how it identifies higher-risk evaluations, approves lowered safeguards, and sets stop conditions.