Home / News / Microsoft Says Only 24 of 8.2 Million Copilot Chats Closely Matched Copyrighted Text
Microsoft Copilot

Microsoft Says Only 24 of 8.2 Million Copilot Chats Closely Matched Copyrighted Text

Sep 6, 20266 min read
Microsoft Says Only 24 of 8.2 Million Copilot Chats Closely Matched Copyrighted Text

News Summary

Microsoft has told a federal court that its Copilot chatbot almost never reproduces meaningful chunks of New York Times articles or other copyrighted works, presenting new usage data as part of a September 4, 2026 (Eastern Time) motion for summary judgment in the sprawling AI copyright litigation brought by news publishers and book authors. The filing, submitted in the U.S. District Court for the Southern District of New York, argues that real-world Copilot usage contradicts claims that the chatbot routinely serves as a free substitute for subscription journalism or copyrighted books.

The Numbers Behind the Claim

Microsoft's legal team worked with an expert hired by the news publisher plaintiffs to analyze 8.2 million Copilot conversation logs. Crucially, this was not a random sample: the logs were deliberately selected because they contained keywords tied to the plaintiffs' websites, making them the conversations most likely to surface copyrighted news content. Even under these favorable conditions for the plaintiffs, the results were sparse.

Fewer than 1 percent of the logs — 59,545 conversations — contained at least 16 words in common with news articles used to ground the underlying AI model's responses. When a stricter threshold was applied by an expert in the related authors' lawsuit, only 24 of the 8.2 million conversations contained at least 30 matching words drawn from books asserted by author plaintiffs. Matches were found in just 10 of the 212 books evaluated, meaning 202 books produced no matching content at all in the sampled logs.

Microsoft also cited findings from an expert working for the Center for Investigative Reporting, one of the news organizations in the consolidated case, who identified 51 examples across the dataset that "substantially overlapped" with CIR's published work. Separately, the plaintiffs' own expert reportedly ran roughly 5.3 million adversarial extraction attempts — deliberately trying to coax the model into reproducing copyrighted text — with fewer than 1 percent of those attempts yielding a 30-word match.

Legal Context: A Consolidated Copyright Fight

The filing is part of In re: OpenAI, Inc. Copyright Infringement Litigation, a multidistrict case overseen by Judge Sidney H. Stein that consolidates claims from The New York Times, the Center for Investigative Reporting, the Authors Guild, and other publishers and writers against OpenAI and Microsoft. The Times originally sued OpenAI and Microsoft in December 2023, arguing that millions of its articles were used without permission to train ChatGPT and Copilot, and that the resulting products could substitute for a Times subscription in some cases. A judge allowed the bulk of the Times' claims to proceed after motions to dismiss in 2025.

On September 4, 2026, Microsoft, OpenAI, and a group of news organizations each filed competing summary judgment motions, asking Judge Stein to resolve core questions of copyright liability and fair use before the case would otherwise proceed to a jury trial. Judge Stein had signed a stipulated sealing order on September 3, 2026, with a September 14, 2026 deadline for parties to request redactions and a September 17, 2026 target for re-filing the briefs publicly with unchallenged portions unredacted. Opposition briefs are due October 5, 2026, with public re-filing of those materials expected around October 15, 2026.

Microsoft's Fair Use Argument

Microsoft's brief leans heavily on the "transformative use" doctrine, arguing that training a large language model to generate natural-language responses — for tasks like coding, drafting, or summarizing — serves a purpose "wholly unlike" the original creative or journalistic purpose of the source material. The company contends that the low reproduction rate in its own data supports this view, stating that "occasional reproduction of source text... hardly undermines the transformative purpose" of the technology.

Microsoft's filing also points to recent rulings in copyright cases involving Meta and Anthropic that found in favor of AI developers on similar fair-use grounds, and the company sought to distance itself from allegations tied to unauthorized data acquisition, stating it did not download books from shadow libraries such as Library Genesis. On the more contested question of whether Bing's search index — which Microsoft has supplied to OpenAI — contained plaintiffs' copyrighted works, Microsoft characterized the evidence as inconclusive.

Market Harm and What Comes Next

Microsoft additionally presented data intended to show no demonstrable decline in book or news sales attributable to generative AI products, along with consumer research suggesting that most readers show little interest in buying AI-generated books even at reduced prices. The company warned that requiring licensing agreements for all AI training data could create what it called insurmountable barriers to innovation.

The Times, CIR, and the Authors Guild reject Microsoft's framing. Their central argument is that Copilot and ChatGPT were built as commercial products capable of substituting for original journalism and literature, competing directly with publishers' subscription and advertising revenue regardless of how often any single response closely mirrors a specific article or book passage. The Authors Guild has previously said such lawsuits "send a clear message to AI companies that authors are taking a strong stand against uses of their works without consent or compensation."

With dueling summary judgment motions now before Judge Stein, the case is moving toward a ruling that could set an influential precedent for how courts nationwide evaluate fair-use defenses in AI copyright disputes, well ahead of any trial.

Microsoft CopilotAI copyright