By Neil Cameron, Lead Analyst, Legal IT Insider
On 24 July, writing about Hugging Face’s breach by an OpenAI bot, I said that the relevant unit of governance was not the model but the whole system through which it could act. I also left one question unanswered:
Nobody has said how many times a frontier model has crossed a boundary its developers believed closed, or how many of those crossings were reported. This one surfaced because the victim found it first.
A week later, Anthropic supplied a partial and uncomfortable answer.
Prompted by OpenAI’s disclosure, Anthropic reviewed 141,006 historical evaluation runs in which its models could have obtained internet access while interacting with a third party evaluation environment that should have been sealed off. It found three incidents, involving six runs and three different Claude models, in which a model within the evaluation environment of Anthropic evaluation partner Irregular reached the open internet and gained unauthorised access to the production infrastructure of three different organisations. The earliest incidents dated to April. Anthropic notified the affected organisations on 27 July. The two it succeeded in reaching had not previously detected the activity. It has not reached the third.
Four months after the earliest incident, an organisation whose production infrastructure was accessed may still not know.
The disclosed incident count is therefore not simply a measure of how often these events occur. It is also a measure of whether anyone conducts the retrospective examination capable of finding them. Anthropic opened its review on 23 July, suspended all cyber evaluations the same day, and had identified all three incidents by the next. The work took a day. Nobody had done it until another company’s failure made it necessary.
The same failure, by a different route
These were not repetitions of the OpenAI incident at a technical level. OpenAI’s models reportedly found and exploited a previously unknown vulnerability in an environment intended to contain them. The Claude models used an internet connection that had been left open by a misconfiguration, arising from a misunderstanding between Anthropic and Irregular.
Anthropic says its models exploited nothing complex: weak passwords, unauthenticated endpoints, an exposed debug page, SQL injection.
One qualification belongs here, because I made the equivalent point about OpenAI last month and it applies with equal force in the other direction. These evaluations ran without the classifiers and monitoring that ship with Anthropic’s generally available models. The safeguards are removed deliberately, in order to measure what the underlying model can do, and Anthropic states that the deployed safeguards would have blocked the behaviour identified. That is the second time in a fortnight that a frontier model has crossed a boundary in a configuration where the controls sold to buyers had been switched off on purpose.
Anthropic characterises the incidents primarily as evaluation-harness and operational failures rather than evidence of a model developing an independent objective. On the disclosed facts that reading is defensible. It is not, however, reassuring in quite the way it may sound.
Anthropic told Claude that it was operating inside a simulation with no internet access. That was not true. When the models encountered real systems, they largely treated them as components of the fictional capture-the-flag exercise they had been instructed to complete. The model did not have to become malicious. Two organisations jointly created an environment whose actual boundaries did not match the boundaries represented to the model, and neither detected the discrepancy before the evaluation began. Once again, no single party governed the composite surface.
That is precisely the structure created when a legal organisation combines an AI provider, an implementation partner, cloud infrastructure, external data sources and access to internal systems. Each participant may control its own component. The agent acts through the combined environment.
What the models did when they worked it out
The three models behaved differently once evidence emerged that their targets were real. Opus 4.7 recognised that it had reached production systems and continued its attack; in two of the four runs it rationalised that the real company must somehow be part of the exercise. Mythos 5 identified in its own reasoning that publishing a package to a real registry would constitute a real-world attack and could not be the intended solution, then reasoned its way back to the conclusion that it remained in a simulation – on the grounds that it did not recognise the certificate authorities securing its connections, and that the systems showed a calendar date of 2026. The third model, an internal research prototype and the most recent of the three, concluded unprompted that its target was real and stopped.
Anthropic reads that ordering as cautious evidence that more capable models respond more appropriately, and qualifies it immediately: three isolated incidents, not a controlled comparison, and further testing required before the conclusion holds. The qualification is the right one, and it should be repeated with the claim. In one case out of three, the control that worked was the model’s own judgment about whether its targets were real. A firm cannot specify or audit that.
An agent damaged a company it did not know existed
Claude Mythos 5 found a document inside the fictional environment: setup instructions telling new developers to install a Python package that did not exist. It built a malicious package under that name and uploaded it to the real PyPI, the public registry for Python software. Anthropic sets out what that took. To register an account the model needed an email address; to obtain one it needed a phone number; it tried and failed to raise funds to pay for a phone number, backtracked, found a free provider and completed the registration. Clever.
The package was available for roughly an hour. It was downloaded and run on 15 real systems. One belonged to a security company whose scanner routinely installs packages in order to examine them for malware. The code executed, extracted credentials, and Claude used those credentials to reach further infrastructure belonging to that company. It had not targeted the company and did not know it existed when it published the package. PyPI’s own security systems removed the package.
That is not merely a containment failure. It is a supply-chain event. An agent pursuing a narrow and authorised objective produced an artefact that damaged a party outside the exercise altogether, through ordinary automated processes.
I have not seen the implication for firms addressed anywhere. An agent working inside an authorised scope, on a matter, can produce something that harms a party with no connection to the engagement. No retainer contemplates that. I doubt many cyber policies do either.
The question for firms
My earlier article asked firms whether their AI-assisted incident-response tooling would function during an attack, where the forensic data would go, and who could override a refusal that obstructed legitimate defensive work.
Anthropic has answered some of what I would have asked next. It retains evaluation transcripts, and it read them. It has committed to continuous monitoring of transcripts for unexpected behaviour and to more rigorous assurance work with the vendors it relies on. METR, an independent evaluation organisation, will conduct a third-party review with access to all transcripts and sampling access to the models, and Anthropic says it will publish a redacted transcript of the PyPI run. Irregular is conducting its own investigation. Measured against the industry’s ordinary conduct, that is a good disclosure.
What none of it establishes is obligation. Every element was voluntary, and the review that produced it began because another company’s failure had become public first. The question that survives is contractual, and it is the only one here a firm can act on at renewal: what in your agreement with an AI provider requires it to tell you if a review it was never obliged to run later finds that your systems were touched?
The OpenAI – Hugging Face incident contained a five-day attribution gap. Anthropic has now disclosed a discovery gap of roughly three months, closed in a day once someone looked.
The first question was whether an autonomous system could cross a boundary believed to be closed. The next is whether anyone will know when it has.
Sources
Anthropic, ‘Investigating three real-world incidents in our cybersecurity evaluations’, 30 July 2026 (anthropic.com/news/investigating-incidents-cybersecurity-evals), updated 3 August 2026, for the review of 141,006 runs, the three incidents, the model-by-model behaviour, the PyPI package, the response measures and the METR review. Irregular (irregular.com) as the evaluation partner. OpenAI, ‘OpenAI and Hugging Face partner to address security incident during model evaluation’, 21 July 2026. Reporting: TechCrunch, 30 July 2026; CNN Business, 30 July 2026; Forbes, 31 July 2026. Earlier analysis: Neil Cameron, ‘The five-day gap: what the OpenAI – Hugging Face incident should tell law firms’, Legal IT Insider, 23 July 2026.







