Court, Explained
U.S. Federal District Courts
Back to docket
N.D. Cal.Procedural orderFiled Dec. 19, 2025

In re Google Generative AI Copyright Litigation

Judge
Van Keulen
Docket
5:23-cv-03440
Court
U.S. District Court · Northern District of California
Pages
2
DiscoveryCivil Procedure
In one sentence

In re Google Generative AI Copyright Litigation: Judge Van Keulen denied plaintiffs’ request to identify works used to train Google’s AI models.

Who this affects

The plaintiffs, including the named plaintiffs and putative class members, and Google LLC. The ruling denied plaintiffs’ request for additional identification-related discovery from Google.

What happened

In In re Google Generative AI Copyright Litigation, plaintiffs sought information identifying the named plaintiffs’ and proposed class members’ copyrighted works that Google LLC used to train its artificial-intelligence models. This was the fourth discovery dispute about that effort.

Plaintiffs asked Google to identify the works or produce evidence sufficient to identify them. The court concluded that only the final training datasets would show which works survived preprocessing, filtering, deduplication, and other steps. It also noted that plaintiffs had already developed a method for identifying works in the training data.

Judge Susan Van Keulen denied the request. She ruled that the proposed relief would require Google to duplicate work plaintiffs had already done, and that the request was redundant and not proportional to the needs of the case.

The detailed version

For law students, journalists, and other readers who want the full reasoning

Case
In re Google Generative AI Copyright Litigation · No. 5:23-cv-03440
Judge
Van Keulen
Date
Dec. 19, 2025

Background

This order addressed the fourth discovery dispute concerning plaintiffs’ efforts to identify the named plaintiffs’ and putative class members’ copyrighted works that Google LLC used to train its artificial-intelligence models. Plaintiffs had previously sought Google’s source code, but the court denied that request. After Google represented that the training datasets—not the source code—would allow the works to be identified, the court approved, with modifications, an inspection protocol for those datasets. Later disputes concerned refining and enforcing that protocol.

Discovery Request

Plaintiffs asked the court to compel Google, by whatever means possible, to produce evidence sufficient to identify, or simply identify, all named plaintiffs’ copyrighted works and all putative class members’ copyrighted works used to train Google’s models.

Plaintiffs pointed to evidence suggesting that Google knew which works were acquired by its Core Data Acquisition Team and ingested through web crawling. They also pointed to records of books ingested into certain other data corpuses.

Court’s Reasoning

The court relied on Google’s position that the only record and means of determining which works were ultimately used to train the models were the training datasets. Information about what Google acquired or placed into source corpuses did not show what survived preprocessing, filtering, deduplication, and construction of the final training datasets. The court therefore concluded that only the training datasets would reveal which putative class members’ works were ultimately used.

The court also reasoned that requiring Google to produce evidence sufficient to identify the works would require an expert to work backward from the training datasets. Plaintiffs acknowledged that they had already developed a methodology that reliably identified class works in Google’s training data. The court concluded that the requested relief would make Google duplicate work plaintiffs had already done, rather than remedy the time and expense plaintiffs incurred developing that methodology.

The court cited the principle that a party generally is not required to create a document when none exists. It also cited Federal Rule of Civil Procedure 33(d), explaining that a defendant generally need not analyze records when the burden of deriving the answer from those records would be the same for the plaintiff.

Disposition

The court found that plaintiffs’ request was redundant and not proportional to the needs of the case. Judge Susan Van Keulen therefore DENIED the request. The order resolved the discovery dispute and did not decide the underlying copyright claims.

The authoritative version

Read the full 2-page opinion on CourtListener, the free public archive maintained by the Free Law Project.

Open opinion PDF →
Summary written with AI assistance. See how summaries are made. Spot something wrong? Tell us.