Lawsuits · OpenAI
In the New York litigation it is undisputed that an OpenAI employee downloaded pirated books from LibGen in 2018. Authors allege much more, including torrenting all of LibGen in 2019 and 2020. The consolidated cases are before Judge Sidney H. Stein.
Each statement is labeled by its source. Findings and orders come from the court; company filings and papers are the company’s own words; allegations are claims not yet proved.
A discovery order describes as undisputed that “in 2018, an OpenAI employee downloaded pirated copies of books from Library Genesis.”
Authors Guild v. OpenAI, S.D.N.Y., Dkt. 782 (Nov. 24, 2025)Class plaintiffs say two employees torrented about 35 TB, all of LibGen’s fiction and non-fiction, between September 2019 and January 2020, and that OpenAI deleted its LibGen data in 2022.
Rule 56.1 statement, MDL Dkt. 1987 (Sept. 17, 2026)Plaintiffs assert that datasets called LibGen1 and LibGen2 were renamed Books1 and Books2; textbook authors allege Books1 and Books2 contain ebooks downloaded from LibGen.
Dkt. 782; Sullivan v. OpenAI, complaint (Aug. 14, 2026)Plaintiffs allege the 2018 download script, written October 10, 2018, took only EPUB and MOBI files. Our reports mark EPUB and MOBI files on every catalog line.
MDL Dkt. 1987 ¶¶ 146–147OpenAI filed the Ninth Circuit’s Doe v. GitHub opinion as supplemental authority for its pending summary judgment motion.
MDL Dkt. 2082 (Oct. 1, 2026)The authors’ and publishers’ suits are consolidated as In re OpenAI, Inc. Copyright Infringement Litigation, No. 1:25-md-03143 (S.D.N.Y.).
MDL docketRegistered works in LibGen as listed at the times OpenAI is found or alleged to have downloaded it.
Registrations matched to dated catalog records in our dataset: our estimate, not a court finding, and each match is a candidate to confirm against the Copyright Office record. Law firms can get these as defendant lists.
Hypothetical maximum scenario (statutory ceiling), list by list. Each list counts works (copies, formats, editions and translations of one book count once) with a registration dated before that copying could have begun (17 U.S.C. § 412). These are legal maximums, not estimates: each figure is what the statute would allow if every work were awarded the $150,000 maximum for willful infringement (§ 504(c)(2)), plus costs and attorney’s fees if a court awards them (§ 505). The statute allows one award per work for all of one infringer’s infringements in the case, or one shared award where infringers are jointly and severally liable, and all parts of a compilation or derivative work count as one work (§ 504(c)(1)). So a company’s lists overlap and must not be added together: a book on two of its lists is still one award, and the figures for different companies are not a combined total.
Counted once across all of OpenAI’s lists: 406,939 works, up to $61.0 billion. Of the 406,939 works on its lists (314,795 on more than one), these have a registration dated before the earliest date OpenAI’s copying of them could have begun. This, not the sum of the lines above, is the per-company figure.
The lists overlap, so the figures are not added together. None of these figures is a prediction or a valuation: ordinary awards run from $750 to $30,000 per work, the maximum requires a finding of willfulness, a work counts once however many copies were made, matches are candidates to confirm, and whether any work qualifies is for a court.
Our snapshots of May 2017, April 2019 and February 2020 fall before and around those dates. The free check lists your titles and counts them in each.
Check a name freeSee a sample reportLast reviewed October 3, 2026. Developments are logged on our news page.