Microsoft says millions of Copilot conversations turned over in discovery rarely reproduced substantial portions of the books and news stories at the center of its copyright fight with publishers and authors.
In court filings submitted Friday, the company said just 24 responses among 8.2 million Copilot conversations contained at least 30 words matching one of the books asserted in the litigation. Only 10 of 212 books reviewed produced any match, according to Microsoft.
The figures are part of Microsoft’s push for summary judgment in the consolidated copyright cases against it and OpenAI, reported by The Verge. Microsoft argues the results bolster its case that using copyrighted works to train AI systems can qualify as fair use and that Copilot generally serves a different purpose from the works used to build it.
But the plaintiffs’ cases are broader than whether Copilot reproduces passages word for word. Publishers and authors also accuse Microsoft and OpenAI of using copyrighted material without permission to train commercial products that compete with their work.
The numbers for news content are larger. Microsoft says an expert hired by the publishers found that 59,545 of the 8.2 million logs contained at least 16 words in common with news content used to ground the AI system. An expert for the Center for Investigative Reporting identified 51 instances that Microsoft characterized as “substantial overlap” with CIR reporting.
The docket filing in Authors Guild v. OpenAI lays out the corresponding figures for books.
Microsoft says the 8.2 million conversations were deliberately selected because they mentioned the publishers’ websites, making them more likely to contain the disputed content. The company argues that the relatively low number of matches therefore strengthens its case.
The New York Times rejects that conclusion. Lead counsel Ian Crosby said evidence uncovered during discovery shows Microsoft and OpenAI used Times journalism to build commercial products that can substitute for the newspaper’s work.
What do 1,000 journalists and PR pros know about AI that you don't? They took AI Quick Start, a 1-hour live class from The Media Copilot. 94% satisfaction. Find out how to work smarter with AI in just 60 minutes. Get 20% off with the code AIPRO: https://mediacopilot.ai/
Microsoft, meanwhile, argues that occasional reproduction “hardly undermines” what it calls the transformative purpose of training large language models. The Trump administration entered the dispute last week, filing a statement of interest supporting OpenAI’s argument that training AI models on copyrighted works can qualify as fair use.
The plaintiffs argue that output overlap is only part of the dispute. Their claims also concern the use of copyrighted material in training and whether the resulting AI products compete with the original works.
Publishers are making related copyright claims elsewhere, including in a European publishers’ suit against Google over AI training and the dozen-plus publisher cases now filed against OpenAI and Microsoft.
Microsoft is also pursuing a different approach outside the courtroom. The company has launched a marketplace to broker AI licensing deals with publishers, creating a system for publishers to make their content available for licensed AI use.
The news and book cases have been consolidated before one judge despite objections from the plaintiffs. Microsoft is now asking the judge to resolve the claims at summary judgment, which could end some or all of them without a trial.
If the court declines to do so, the litigation continues, along with the disagreement over what those 8.2 million Copilot conversations actually prove.







