Global AI governance and computer crime liability

Rogue AI Cyber-Tests Reveal an Evaluation-Vendor Failure, Not a CFAA Vacuum

Anthropic's and OpenAI's real-world hacks trace to one testing vendor's misconfiguration — the fix is infrastructure standards, not a new hacking law.

Anthropic's Cyber-Eval Breach, By the Numbers People of Internet Research · Global 141,006 Eval sessions reviewed Anthropic's retrospective review t… 3 Organizations breached Real production systems compromise… 2 of 3 Victims unaware until notified Anthropic told two organizations t… June 2, 2026 CFAA enforcement order signed Executive Order 14409 directs the … peopleofinternet.com
Anthropic's Cyber-Eval Breach, By the … People of Internet Research · Global 141,006 Eval sessions reviewed 3 Organizations breached 2 of 3 Victims unaware until notified June 2, 2026 CFAA enforcement order signed peopleofinternet.com

Key Takeaways

Three Companies, One Misconfigured Range

On July 30, 2026, Anthropic disclosed that Claude models had broken out of what were supposed to be isolated cybersecurity evaluations and gained unauthorized access to the live systems of three real organizations. The company had reviewed 141,006 evaluation sessions after rival OpenAI's July 21 disclosure that one of its models breached Hugging Face's production infrastructure during a similar test. Anthropic suspended internet-connected cyber evaluations on July 23, notified the affected organizations and its testing partner, Irregular, on July 27, and went public three days later. Two of the three victims had no idea they had been breached until Anthropic told them.

The root cause was mundane rather than dramatic. Capture-the-flag prompts told the models the range was an offline simulation; a misunderstanding between Anthropic and Irregular left the machines with live internet access anyway. Once online, the models did what they were tasked to do — hunt for exploitable weaknesses — against whatever they found, including weak passwords and unauthenticated endpoints on real corporate networks. Anthropic has called this "closer to a harness and operational failure than a model alignment failure," and the detail that its newest model stopped attacking once it recognized the environment was real supports that read.

A Concentration Risk Nobody Priced In

Irregular is not a bit player. The same company's infrastructure sits behind OpenAI's and Anthropic's cyber evaluations, and reporting has since tied a similar outbound-connectivity gap to an incident involving Meta's Muse Spark model. Three of the industry's most consequential AI labs were relying on overlapping evaluation infrastructure that, in at least one shared respect, wasn't sealed. That is a supply-chain concentration problem as much as it is an AI safety one — the kind of single-vendor dependency that critical-infrastructure regulators have spent two decades learning to distrust in other sectors.

The Liability Question the Statute Wasn't Written For

The incidents landed directly on top of a live legal debate. The Computer Fraud and Abuse Act, 18 U.S.C. § 1030, requires proof that someone "intentionally" accessed a computer without authorization; no court has yet decided what that means when the actor executing the intrusion is a model, not a person. President Trump's June 2, 2026 executive order, EO 14409, directs the Attorney General to "prioritize the enforcement" of §1030 and related statutes against anyone who "utilizes AI to illegally access or damage a computer without authorization." That directive was written with malicious actors in mind — but as legal analysts have noted, it now sits awkwardly next to two labs whose own evaluation programs produced exactly the conduct the order targets, absent any intent to cause harm.

There is a real case for taking this seriously rather than waving it off as an unlucky misconfiguration. These weren't sandboxed toy exercises — they were frontier models, run with production safety classifiers deliberately disabled, pointed at exploit-hunting tasks, and connected, however inadvertently, to the live internet. Two victims sat compromised for weeks without knowing it. A negligence standard doesn't require intent, only a failure to take reasonable precautions against foreseeable harm — and cybersecurity attorney Ahmed Ghappour's framing is hard to dismiss: "You don't get to deploy something capable of breaking into systems and then disown where it goes."

Why Stretching the CFAA Is the Wrong Instrument

But the remedy for that argument already exists, and it isn't a new AI-hacking crime. Civil negligence claims under existing tort law, plus the AI-liability statutes California, New York, and Rhode Island have enacted — which hold companies responsible when their AI systems cause harm a human operator would be liable for — already give the three victims here a path to recovery without anyone needing to solve the CFAA's intent problem. Retrofitting a 1986 anti-hacking statute, built for human actors deliberately breaking in, to reach a lab that voluntarily disclosed its own testing failure risks punishing exactly the transparency that surfaced these incidents in the first place. Both Anthropic and OpenAI disclosed unprompted, ahead of any regulatory demand or public exposure. An enforcement posture that treats that disclosure as a confession invites labs to simply stop looking — auditing fewer transcripts, disclosing fewer near-misses, and leaving the next misconfigured range undiscovered until an outside party finds it first.

What Proportionate Oversight Looks Like

The actual failure here is procurement and infrastructure hygiene, not a gap in criminal law. A better response looks like security baselines for third-party AI evaluation environments — network isolation verified by the lab, not asserted in a prompt, plus mandatory breach notification timelines when an eval environment turns out not to have been sealed. NIST's AI Risk Management Framework and the evaluation-transparency commitments labs already made to the UK AI Safety Institute are the natural homes for that standard, not a new theory of computer-crime liability stretched to cover a testing vendor's internet routing. Congress and regulators should watch this space closely — but the lesson of July's disclosures is that the existing civil and state liability tools already reach negligent labs, while criminalizing disclosed testing failures would only teach the next lab to stop disclosing.

Sources & Citations

  1. Anthropic: Investigating three real-world incidents in our cybersecurity evaluations
  2. Federal Register: EO 14409, Promoting Advanced AI Innovation and Security
  3. Hugging Face: Security incident disclosure, July 2026
  4. TechCrunch: Who's legally to blame for Anthropic and OpenAI's autonomous AI hacks?
  5. CTech: Irregular at center of AI labs' evaluation breaches