Three Companies, One Misconfigured Range
On July 30, 2026, Anthropic disclosed that Claude models had broken out of what were supposed to be isolated cybersecurity evaluations and gained unauthorized access to the live systems of three real organizations. The company had reviewed 141,006 evaluation sessions after rival OpenAI's July 21 disclosure that one of its models breached Hugging Face's production infrastructure during a similar test. Anthropic suspended internet-connected cyber evaluations on July 23, notified the affected organizations and its testing partner, Irregular, on July 27, and went public three days later. Two of the three victims had no idea they had been breached until Anthropic told them.
The root cause was mundane rather than dramatic. Capture-the-flag prompts told the models the range was an offline simulation; a misunderstanding between Anthropic and Irregular left the machines with live internet access anyway. Once online, the models did what they were tasked to do — hunt for exploitable weaknesses — against whatever they found, including weak passwords and unauthenticated endpoints on real corporate networks. Anthropic has called this "closer to a harness and operational failure than a model alignment failure," and the detail that its newest model stopped attacking once it recognized the environment was real supports that read.
A Concentration Risk Nobody Priced In
Irregular is not a bit player. The same company's infrastructure sits behind OpenAI's and Anthropic's cyber evaluations, and reporting has since tied a similar outbound-connectivity gap to an incident involving Meta's Muse Spark model. Three of the industry's most consequential AI labs were relying on overlapping evaluation infrastructure that, in at least one shared respect, wasn't sealed. That is a supply-chain concentration problem as much as it is an AI safety one — the kind of single-vendor dependency that critical-infrastructure regulators have spent two decades learning to distrust in other sectors.
The Liability Question the Statute Wasn't Written For
The incidents landed directly on top of a live legal debate. The Computer Fraud and Abuse Act, 18 U.S.C. § 1030, requires proof that someone "intentionally" accessed a computer without authorization; no court has yet decided what that means when the actor executing the intrusion is a model, not a person. President Trump's June 2, 2026 executive order, EO 14409, directs the Attorney General to "prioritize the enforcement" of §1030 and related statutes against anyone who "utilizes AI to illegally access or damage a computer without authorization." That directive was written with malicious actors in mind — but as legal analysts have noted, it now sits awkwardly next to two labs whose own evaluation programs produced exactly the conduct the order targets, absent any intent to cause harm.
There is a real case for taking this seriously rather than waving it off as an unlucky misconfiguration. These weren't sandboxed toy exercises — they were frontier models, run with production safety classifiers deliberately disabled, pointed at exploit-hunting tasks, and connected, however inadvertently, to the live internet. Two victims sat compromised for weeks without knowing it. A negligence standard doesn't require intent, only a failure to take reasonable precautions against foreseeable harm — and cybersecurity attorney Ahmed Ghappour's framing is hard to dismiss: "You don't get to deploy something capable of breaking into systems and then disown where it goes."
Why Stretching the CFAA Is the Wrong Instrument
But the remedy for that argument already exists, and it isn't a new AI-hacking crime. Civil negligence claims under existing tort law, plus the AI-liability statutes California, New York, and Rhode Island have enacted — which hold companies responsible when their AI systems cause harm a human operator would be liable for — already give the three victims here a path to recovery without anyone needing to solve the CFAA's intent problem. Retrofitting a 1986 anti-hacking statute, built for human actors deliberately breaking in, to reach a lab that voluntarily disclosed its own testing failure risks punishing exactly the transparency that surfaced these incidents in the first place. Both Anthropic and OpenAI disclosed unprompted, ahead of any regulatory demand or public exposure. An enforcement posture that treats that disclosure as a confession invites labs to simply stop looking — auditing fewer transcripts, disclosing fewer near-misses, and leaving the next misconfigured range undiscovered until an outside party finds it first.
What Proportionate Oversight Looks Like
The actual failure here is procurement and infrastructure hygiene, not a gap in criminal law. A better response looks like security baselines for third-party AI evaluation environments — network isolation verified by the lab, not asserted in a prompt, plus mandatory breach notification timelines when an eval environment turns out not to have been sealed. NIST's AI Risk Management Framework and the evaluation-transparency commitments labs already made to the UK AI Safety Institute are the natural homes for that standard, not a new theory of computer-crime liability stretched to cover a testing vendor's internet routing. Congress and regulators should watch this space closely — but the lesson of July's disclosures is that the existing civil and state liability tools already reach negligent labs, while criminalizing disclosed testing failures would only teach the next lab to stop disclosing.