Doe v. GitHub, Decided 16 September 2026: Why the Ninth Circuit Held That Copilot Creates New Works Under DMCA Section 1202(b), and How the Training-Data Theory Was Forfeited Before It Was Argued
A magazine page carries its photographer's name in the gutter, the narrow margin beside the image where the credit is printed in small type. Scan the page, crop the margin, post the picture, and a court can look at the two versions side by side: the same composition, the same crop, and one line of text gone. That comparison is the paradigm the Digital Millennium Copyright Act had in mind in 1998 when it made it unlawful to "remove or alter" copyright management information, and it is the example the Ninth Circuit reached for on 16 September 2026 when it explained why the same statute does not reach a coding model that writes a function resembling one it was trained on.
The case is Doe v. GitHub, Inc., No. 24-7700, an interlocutory appeal from the Northern District of California, argued in San Francisco on 11 February 2026 before Judges Sidney Thomas and Eric Miller and District Judge Stanley Blumenfeld, sitting by designation. Judge Miller wrote for a unanimous panel. The disposition is one word, affirmed, and the reasoning runs eighteen pages. What follows reconstructs how the claim was built, how half of it disappeared before the appeal, and what the half that reached the court was held to require.
Practical takeaway. After 16 September 2026, a DMCA claim against a generative model in the Ninth Circuit needs facts showing that copyright management information was taken off a copy of an existing work. Output that is similar, even substantially similar, to training data is a question for ordinary copyright infringement, which the panel left open, and for the contract claims still pending before Judge Tigar. The training-stage theory was never decided; it was forfeited on this record, and a differently pleaded case could raise it.
What the Programmers Alleged in Case 4:22-cv-06823, and What Survived Two Rounds of Dismissal
The plaintiffs are programmers who published code on GitHub under open-source licenses, one of whose most common conditions is attribution: a copy of the license, with the author's name and copyright notice, must travel with any copy or derivative of the code. Their putative class action against GitHub, Microsoft and a long list of OpenAI entities alleged that Copilot and Codex, large language models trained on "all available public GitHub repositories," sometimes emit "identical copies of code Copilot was trained on" without the license text, notice or author's name that accompanied the original. After two rounds of dismissals and amendments before Judge Jon S. Tigar, three claims were left: one under 17 U.S.C. § 1202(b) and two for breach of contract.
Judge Tigar dismissed the DMCA claim, first with leave to amend and then with prejudice, on the reasoning that § 1202(b) claims "require that copies be 'identical'" and that every example in the complaint involved a "modified format," a "variation" or the "functional equivalent" of the licensed code. He denied the motion to dismiss the contract claims and certified one question for appeal under 28 U.S.C. § 1292(b): whether §§ 1202(b)(1) and (b)(3) impose an identicality requirement. The amicus roster shows how much was thought to ride on the answer: the App Association; the Chamber of Progress with the Computer & Communications Industry Association; the Authors Guild, the Association of American Publishers, the News/Media Alliance and the STM publishers; the Electronic Frontier Foundation with Public Knowledge; the Authors Alliance; and a group of intellectual property professors represented by Berkeley's Samuelson Clinic.
How the Input Theory Was Forfeited: "Perhaps It Doesn't" and "The Complaint Is Not About Training"
The plaintiffs had two theories. The "input" theory said the defendants violated § 1202(b)(1) at the training stage by stripping copyright management information from class members' code before feeding it into the model. The "output" theory said that when the model returns memorized training data to a user without the original notices, it removes or alters that information on a copy of the work. Only the second was decided.
The panel's account of why is a lesson in how litigation positions harden. The operative complaint could be read to plead an input theory; it alleges that the defendants removed CMI "by incorporating it into Copilot with its CMI removed." But at the hearing on the first motion to dismiss, when Judge Tigar asked whether copying training data into Copilot violated the licenses' attribution requirement, plaintiffs' counsel answered, "Perhaps it doesn't." The judge replied that the "complaint is not about training. It just isn't," and his subsequent order stated that the plaintiffs "do not allege they were injured by Defendants' use of licensed code as training data." Nothing in the later briefing told the court it had misunderstood. On appeal the plaintiffs pointed to one sentence about "the mere removal of CMI from digital copies" and to an OpenAI filing, which the panel read as saying the opposite: that the second amended complaint "admits that no pre-distribution [CMI] removal occurred." The theory was held forfeited, and the opinion says nothing about whether it would have succeeded.
Why the Panel Found Standing: Memorization Research and GitHub's Own 150-Character Duplicate Filter
The defendants argued that even if Copilot copies code without notices, the chance that it would reproduce these plaintiffs' code, rather than some other contributor's, was conjecture. The panel disagreed, and its reasons are worth reading by anyone who builds or buys a coding assistant. The complaint cites academic research that large language models "emit the memorized training data verbatim" and that the tendency "will likely get worse as models continue to scale." It gives examples of Copilot reproducing portions of the named plaintiffs' own code verbatim. And it points to a feature GitHub itself ships: a duplicate-detection setting that lets users block suggestions matching public code, specifically verbatim snippets of 150 characters or more, which the court called "some evidence that Copilot can and does emit literally identical copies of code."
That is a notable use of a vendor's safety control. The filter exists to reassure enterprise customers; here it helped establish that the risk it guards against is real enough for Article III. The panel added that at summary judgment the plaintiffs would need evidence of a "substantial risk," and declined to say whether the material cited in the complaint would meet that bar.
Why "Remove or Alter" Requires an Existing Copy: Gutter Credits, Search Engines and the $25,000-Per-Violation Problem
On the merits the panel started, as it put it, "where we always do: with the text of the statute." "Remove" means to get rid of; "alter" means to cause to become different. Both verbs, in context, "imply taking an affirmative act with respect to CMI connected to a work that already exists." The statutory definition points the same way: CMI is information conveyed "in connection with copies" of a work, and a copy is a material object in which the work is fixed. Defacing a title page, deleting the metadata that accompanies a file, cropping the gutter credit from a reprinted photograph: each takes something off a copy. Someone who creates a new work and leaves the credit out has removed nothing. The panel cited Falkner v. General Motors, where a photographer who shot a mural from an angle that hid the artist's signature had not altered CMI, and the Fifth Circuit's Kipp Flores Architects decision of 21 August 2026 for the same reading.
The panel then corrected the district court's vocabulary. "Identicality," it wrote, "is something of a misnomer," a gloss on the words remove, alter and copies rather than an independent element. If two works are otherwise identical and one lacks the notice, a factfinder may infer removal, as the Ninth Circuit did in Friedman v. Live Nation. Minor cosmetic changes will not save a defendant who reproduces a work substantially and strips the credit; the opinion cites Real World Media v. Daily Caller for the point that disseminating 99 percent of a copied work is no escape, and Judge Stein's 2025 ruling in New York Times Co. v. Microsoft for the same principle. What sank the plaintiffs was their own description of the machine. Copilot, the complaint says, infers "statistical patterns governing the structure of code" and predicts "the most likely solution to a given prompt." That describes learning from existing works and generating new ones, and the panel drew the contrast explicitly: a search engine retrieves and displays stored copies, and "if Copilot functioned like a search engine and produced outputs that were identical to plaintiffs' code but did not contain CMI, then plaintiffs might have a stronger claim."
The final paragraph explains the stakes as the court saw them. Copilot's output may in some cases be substantially similar to existing code, and the panel expressed "no view on whether that similarity would allow plaintiffs to assert a claim for copyright infringement." But many infringement cases involve a similar work made without attribution, and if that alone violated § 1202(b), the DMCA "would supplant traditional copyright protections" and expose defendants to its enhanced statutory damages, up to $25,000 per violation against the ordinary $30,000 cap per work. The court declined "to transform run-of-the-mill copyright-infringement claims into DMCA claims."
What Remains Open After 16 September 2026: Substantial Similarity, Two Contract Claims, and the EU's Article 53 Rule on the Input Side
Three things survive the ruling. The two contract claims, resting on the licenses' attribution terms, continue before Judge Tigar. Ordinary copyright infringement for substantially similar output is expressly undecided. And the input theory, the one that would have asked whether stripping notices from code before training is itself a removal, is undecided too, because it was forfeited rather than rejected. A plaintiff who pleads training-stage removal from the first complaint onward, and says so at the first hearing, presents a question this opinion does not answer.
For a European reader the contrast with the EU AI Act is instructive. Article 53(1)(c) and (d), applicable to general-purpose model providers since 2 August 2025, put the obligation on the input side by regulation: a copyright-compliance policy that honors rights reservations under the 2019 Digital Single Market directive, and a public summary of training content. One transition applies: under Article 111(3), providers of general-purpose models placed on the market before 2 August 2025 have until 2 August 2027 to comply. Where the US appellate court has declined to reach the training stage because the plaintiffs let it go, the European legislature reached it first and did so without a lawsuit. The two regimes now diverge at exactly the point the Ninth Circuit did not decide.
The humane point is about attribution as the currency of a commons. Open-source developers publish for reputation and reciprocity; the license condition that a name travels with the code is how that economy works. For these plaintiffs, on the output theory they pleaded, enforcement of that condition now runs through contract and, perhaps, infringement, and through a product setting that blocks 150-character matches; a § 1202(b) claim remains available where removal from an existing copy is adequately alleged. Whether the working developer whose function reappears unattributed in a colleague's editor thinks that is enough is a fair question, and the opinion does not pretend to answer it.
How Quentir Reads It
This is the second time in six weeks the Ninth Circuit has decided what an AI system is for the purposes of a specific statute, and the two rulings arrived in different procedural settings. On 4 August 2026, in Amazon.com Services v. Perplexity AI, No. 26-1444, which we read in our 4 September analysis of four documents that named a requirement before its test existed, the court vacated a preliminary injunction: on the record then before it, and judging likelihood of success, it treated Perplexity's browser agent as a tool and attributed the access under the Computer Fraud and Abuse Act to the user. On 16 September, in Doe v. GitHub, the court ruled on a motion to dismiss, taking the complaint's allegations as true, and treated Copilot as something that learns and creates. The first is provisional and evidence-based; the second is final as to this complaint and rests on the plaintiffs' own description of the model. Both opinions start from Van Buren v. United States and its insistence on statutory text, and both classify the software by its function before applying the statute. A functional taxonomy of AI systems is being built in the courts one statute and one record at a time.
That has a practical consequence for anyone drafting a complaint, a contract or a procurement clause about a model. The description of how the system works is now doing legal work. The programmers lost the DMCA claim on their own account of a "complex probabilistic process," and lost the input theory on a hearing exchange. A buyer of coding assistants should read the ruling the other way round: the duplicate-detection filter that helped the plaintiffs establish standing is also the control a vendor can be asked to switch on by default, and its 150-character threshold is a number a procurement document can name.
Quentir has followed this line of decisions as it developed, from the CFAA ruling in August to the DMCA ruling in September, alongside the EU's input-side obligations. The free coverage lives on the blog; the All-access membership is the archive argument for the paid work beside it: every Signature Brief and Signature Report as it publishes, and the full archive, under one organization-wide license for the length of the subscription.
The unresolved question after 16 September 2026 is whether stripping licenses and notices from code before training is itself a removal of copyright management information under § 1202(b)(1). This panel did not decide it, because on this record the theory was not preserved.
Sources: United States Court of Appeals for the Ninth Circuit, J. Doe 1 et al. v. GitHub, Inc. et al., No. 24-7700, opinion by Judge Miller (filed 16 September 2026; argued 11 February 2026; appeal from N.D. Cal. No. 4:22-cv-06823-JST), for every quotation, citation and procedural fact about that case above; United States Court of Appeals for the Ninth Circuit, Amazon.com Services, LLC v. Perplexity AI, Inc., No. 26-1444 (filed 4 August 2026), for the preliminary-injunction disposition and the tool reading under the Computer Fraud and Abuse Act; 17 U.S.C. § 1202 and 17 U.S.C. § 1203 (Legal Information Institute) for the statutory text and damages figures; Courthouse News Service, "Coders lose appeal in copyright fight against AI tools" (16 September 2026), for the same-day report; Regulation (EU) 2024/1689 (AI Act), Article 53(1)(c) and (d), Article 111(3) and Article 113, for the training-content summary and copyright-policy obligations applicable from 2 August 2025 and the 2 August 2027 transition for models placed on the market before that date. Public sources checked 19 September 2026.
Published intelligence, built to inform your own decisions. Published: September 19, 2026.