US AI and copyright

The Ninth Circuit's Copilot Ruling Keeps DMCA Section 1202 to Removal, Not Absence

In Doe v. GitHub, the Ninth Circuit held that AI code output lacking attribution is a new work, not a copy stripped of CMI. Copyright claims are left intact.

Doe v. GitHub at a Glance People of Internet Research · US $2,500 Statutory damages floor Minimum per Section 1202 violation… 150+ chars GitHub duplicate-filter thre… Verbatim snippet length Copilot's … 3 Claims left after dismissals One DMCA claim and two contract cl… peopleofinternet.com
Doe v. GitHub at a Glance People of Internet Research · US $2,500 Statutory damages floor 150+ chars GitHub duplicate-filter… 3 Claims left after dismissals peopleofinternet.com

Key Takeaways

On September 16, 2026, the Ninth Circuit decided Doe v. GitHub, No. 24-7700, and affirmed the dismissal of the plaintiffs' DMCA claim against GitHub, Microsoft and OpenAI. Judge Eric Miller wrote for a panel that also included Judge Sidney Thomas and District Judge Stanley Blumenfeld, Jr. The holding is narrow, but it matters. Missing attribution in AI output is not, on its own, the "removal" of copyright management information under Section 1202(b).

What the plaintiffs argued

The plaintiffs are programmers who published open-source code on GitHub. Copilot is a paid tool that GitHub and OpenAI built on Codex, and it was trained on billions of lines of public code. The plaintiffs alleged that it sometimes emits "essentially verbatim" copies of their code without the author names, copyright notices and license terms that came with the original. They said this violates 17 U.S.C. § 1202(b). By the time of the appeal, two rounds of dismissals had cut the complaint to one DMCA claim and two breach-of-contract claims, according to the opinion. The district court dismissed the DMCA claim on the theory that Section 1202(b) requires "identical" copies. It then certified that question for interlocutory appeal.

The strongest case for the plaintiffs

The plaintiffs' position deserves a fair statement. Open-source licenses are conditional. Attribution is often the price of free reuse, and a tool that emits licensed code without the license text guts that bargain at industrial scale. The complaint also alleges that memorization of training data "will likely get worse as models continue to scale." It notes that GitHub itself built a duplicate-detection filter that blocks suggestions matching public code of 150 characters or more. That is some evidence that verbatim output happens. If Congress meant Section 1202 to protect attribution metadata, the argument goes, it should not matter whether a machine or a person did the stripping.

Why the court said no

The panel started with the statute's text. Section 1202(b)(1) bars anyone from "intentionally remov[ing] or alter[ing]" CMI. The court read "remove" as "to get rid of" and "alter" as "to cause to become different". Both verbs, it said, imply an affirmative act on CMI "connected to a work that already exists." In the opinion's words, one who creates a new work and fails to include CMI cannot be said to have "removed" or "altered" anything. The statutory definition points the same way. CMI is information "conveyed in connection with copies" of a work, so it lives on the material object, and the plaintiffs alleged no removal from such a copy. The court described Copilot as a tool that "does not look up and reproduce stored work but rather creates new work." That new work "may or may not infringe plaintiffs' copyrights," the court said, but it "cannot reasonably be described as a copy of that code from which CMI has been 'removed' or 'altered.'"

The court did not adopt the district court's "identicality" rule. It called identicality "something of a misnomer" and said the DMCA does not require literal identity. Near-copies with cosmetic changes can still support an inference of removal. It cited Friedman v. Live Nation (9th Cir. 2016) and Real World Media v. Daily Caller (D.D.C. 2024) on this point. The plaintiffs simply had not alleged that any CMI was stripped from a copy of their work.

The court also held that the plaintiffs had Article III standing, because they plausibly alleged a substantial risk of injury. It treated their "input" theory, that CMI was stripped before training, as forfeited. Plaintiffs' counsel had told the district court the training theory might not be their claim, answering "Perhaps it doesn't."

What the ruling does not decide

This is not a general ruling that AI training or output is lawful. The opinion says the output may still infringe, and that question is left for the copyright claims. The breach-of-contract claims in this case also remain pending in the district court. Nor does the ruling shield a system that actually strips CMI from a copy of a specific work. Under Section 1203, statutory damages for each Section 1202 violation run from $2,500 to $25,000. A per-output metadata theory would have multiplied that exposure across every generation ever produced, and that is the outcome the ruling forecloses.

Why the pro-innovation reading is right

Section 1202 is a narrow anti-tampering provision. As EFF noted, an expansive reading would have created new liability for remixes, classroom adaptations, reverse engineering and search indexing, because all of those produce works without the original's metadata. Attaching statutory damages to the absence of a label would turn a labeling rule into a liability engine for anything transformative.

The better path for rightsholders is the one the court left open. Where an AI system reproduces protected expression, copyright infringement and license-breach claims test whether that reproduction is actually unlawful, and they do so on evidence. That is how the Supreme Court handled a comparable panic in Sony v. Universal City Studios (1984). It found that a technology capable of substantial noninfringing uses was not contributory infringement, and that the studios' evidence of market harm was speculative. The Ninth Circuit has similarly kept a new technology from being condemned through a statute written for a different problem.

Developers should not read this as a free pass. Memorization is a real technical phenomenon. Filters like GitHub's 150-character duplicate detector, and attribution tooling generally, are the proportionate answer that both protects licensors and keeps tools useful. Policymakers who want attribution mandated for AI output should legislate that duty explicitly rather than stretch 1998 statutory language. Section 1202 was drafted for scrubbed metadata, not for generative models.

Sources & Citations

  1. Ninth Circuit opinion, Doe v. GitHub (No. 24-7700, Sept. 16, 2026)
  2. 17 U.S.C. § 1203 (civil remedies, statutory damages)
  3. Sony Corp. of America v. Universal City Studios (1984)
  4. EFF: Victory! Appeals Court Rejects Expansive New Copyright Claim