The fight in the AI copying cases has moved from whether a company copied a pirate library to which works were in it, who owns them, and when. Motions to dismiss now test ownership, timeliness and distribution work by work. PayAuthors supplies the work-level facts from the libraries' own published catalogs and the U.S. Copyright Office's records.
What the record already says
Courts and the defendants themselves have supplied the collection-level facts. Each statement below is labeled by its source.
| Statement | Source | Kind |
|---|---|---|
| Anthropic's co-founder "used the infamous BitTorrent protocol" to copy LibGen in June 2021. | Bartz v. Anthropic, N.D. Cal., ECF No. 244 at 2 | Court order |
| Anthropic compared PiLiMi with its LibGen copies and torrented only the copies "that [we]re not in Lib[G]en." | Bartz, ECF No. 244 at 3 | Court order |
| Every work on the settlement Works List was downloaded by Anthropic by August 10, 2022. | Bartz, ECF No. 418 at 49 | Stipulation |
| Contributory liability requires intent, shown by inducement or by a service tailored to infringement. | Cox Communications v. Sony Music (U.S. 2026) | Supreme Court |
| Meta did not dispute, at the pleading stage, that uploading and downloading are infringing acts. | Strike 3 Holdings v. Meta, N.D. Cal., ECF No. 61 at 7 | Court order |
What remains to prove is the work-level link: that a given owner's book was in the specific files and torrents taken.
What we produce for each work
One line from a live report, with the five facts a complaint exhibit needs. Every report and defendant list is built from lines like this.
The Lost Night
- 2LibGen fiction ID 2157462 · epub · MD5
8b2a0eab82719cfd5ab89c3e4fd6e3de - 3torrent
f_2157000· file 12 MiB, spans 4 pieces of 4.0 MiB; largest share in one piece 33%
- Publication: published Feb 26, 2019
- Authorship: Bartz, Andrea — author
- Claimants: Andrea Bartz Inc., (organization)
- Recorded documents: none in the Office’s bulk data through Jun 26, 2025
- 1Catalog entry: the record as LibGen listed it, the date it was added, and which of 25 dated catalogs list it.
- 2Exact file: catalog number, format and MD5, the fingerprint of one specific file.
- 3Torrent: the named torrent that carried the file, and how the file sits in its pieces.
- 4Registration: the Copyright Office record, with the claimant (the rights holder named on it), dates and any recorded transfers.
- 5§ 412 timing against each defendant’s documented or alleged download date.
Each work is also screened against the Bartz Works List criteria.
Scale
Distribution claims: a whole book often fits in one torrent piece
Defendants argue that BitTorrent uploads carried only fragments. The torrents' own file lists show that most files fit inside a single piece, so any upload of that piece carried the whole file. They do not show that a defendant uploaded a particular piece; that has to come from the defendant's own records or other evidence. We read them from public archived copies; no swarm is joined and no book file is read.
Pieces are mostly 4 MiB (LibGen 1 to 8 MiB). Figures cover located files in the current dataset. An expert can recheck every figure with a short standard-library script we supply.
Defendant lists
Ready lists of registrations by defendant and theory, each with catalog IDs, MD5s, torrents, piece data and § 412 timing. The basis column says what the defendant-side link rests on.
Read the ceiling column as a hypothetical maximum scenario, not an estimate. It is what the statute would allow if every listed work were awarded $150,000 for willful infringement, plus costs and attorney’s fees if a court awards them. The statute allows one award per work for all of one infringer’s infringements in the case, or one shared award where infringers are jointly and severally liable, and all parts of a compilation or derivative work count as one work (17 U.S.C. § 504(c)(1)): a book on two of a company’s lists is still one award, so lists for the same company must not be added together, and different companies’ figures are not a combined total. It is not a prediction, valuation or damages figure: ordinary awards run from $750 to $30,000 per work, the lists overlap, and whether any work qualifies is for a court.
| List | Works | Hypothetical maximum statutory ceiling, not an estimate | Basis |
|---|---|---|---|
| Anthropic: Books3 works not on the Works List by our screen | 4,515 | $677M | Bartz record |
| Anthropic: LibGen works not on the Works List by our screen | 5,461 | $819M | Bartz record |
| Anthropic: PiLiMi works not on the Works List by our screen | 1,629 | $244M | Bartz record |
| Anthropic: works opted out of the Bartz settlement (their owners kept the right to sue) | 1,802 works | — | Court order (Bartz final approval) |
| Of these, already in a lawsuit against Anthropic | about 620 | — | Our estimate (see note below) |
| Not yet in any lawsuit we found | about 1,180 | — | Our estimate (see note below) |
| Meta: LibGen | 493,479 | $74.0B | Court opinion (Kadrey) |
| Meta: Z-Library records added Aug. 2022 to Aug. 2023, via Anna's Archive (early 2024) | 221,052 | $33.2B | Court opinion (Kadrey); Meta's answer in Elsevier |
| DeepSeek: same Z-Library records, via Anna's Archive (by March 2024) | 209,902 | $31.5B | Company paper (DeepSeek-VL) |
| OpenAI: LibGen, 2018 | 315,229 | $47.3B | Court order; scope alleged |
| OpenAI: LibGen torrenting, 2019 to 2020 | 406,913 | $61.0B | Allegation |
| NVIDIA: same Z-Library records, via Anna's Archive (August 2023) | 217,965 | $32.7B | Allegation; NVIDIA's answer admits downloading from Anna's Archive's datasets page, undated |
Works = books with a registration dated before the copying could have begun (17 U.S.C. § 412), counting copies, formats, editions and translations of one book once. Statutory ceiling = works × $150,000, the maximum for willful infringement (17 U.S.C. § 504(c)(2)), plus costs and attorney’s fees if a court awards them (§ 505). It is the legal maximum, not a prediction: ordinary awards run from $750 to $30,000 per work, the maximum requires a finding of willfulness, lists overlap and are not additive, and each match is a candidate to confirm. Opt-out works are counted by the court’s list, without our timing check.
One award per work: each company’s lists combined
A book on two of a company’s lists is one work and one award. Here each company’s lists are combined and each work is counted once. Because statutory damages depend on registration before the infringement commenced (17 U.S.C. § 412), a work is counted only if a registration predates the earliest date that company’s copying of it could have begun across all its lists.
| Company | Works on its lists | On two or more lists | Counted (registered before its earliest copying) | Hypothetical maximum statutory ceiling, not an estimate |
|---|---|---|---|---|
| Anthropic Books3, LibGen and PiLiMi works not on the Works List | 9,700 | 1,635 | 9,584 | $1.4B |
| Meta LibGen, Anna’s Archive and Books3 | 618,350 | 145,360 | 616,230 | $92.4B |
| NVIDIA Books3 and Anna’s Archive (allegations) | 267,561 | 29,419 | 266,677 | $40.0B |
| OpenAI LibGen 2018 (court order; scope alleged) and LibGen torrenting 2019–2020 (alleged) | 406,939 | 314,795 | 406,939 | $61.0B |
About the opt-out figures
1,802 is the number of works on the court’s opt-out list, from the 350 owners who opted out on time (Bartz ECF No. 680 at 6; list at ECF No. 619-3, Ex. J). The court also allowed 2 late opt-outs; their works are not on that list (ECF No. 680 at 12–14). To split the list, we compared it with the works named in the seven suits opt-outs have filed against Anthropic as of October 1, 2026. For two of those suits we used the same plaintiffs’ exhibits from their suits against Meta. A work can have several owners, and a suit we have not found is not counted, so both figures are estimates, not court findings.
Ongoing monitoring
We check for new filings in the AI copyright cases every hour. New court rulings, company statements and allegations are logged with their sources and graded; documented download dates go into the report timeline and the defendant lists after review, and newly archived catalogs are added in dataset releases. The latest items are on our News page. For firm clients we can flag developments that touch their clients' works.
Ways to work with us
Paid pilot: $1,000, fixed scope
Up to 50 distinct works, chosen from our defendant lists or by registration numbers you give us. For each work: the registration and its Copyright Office record, every catalog line with its ID, MD5, torrent and piece data, the registration-date screen against each company, and a note on any edition or same-name risk. Every line is reviewed by PayAuthors against the Office record before delivery, and you get one correction pass. Delivered as a spreadsheet and PDF exhibits with a method note, in five business days. The fee is credited toward a larger engagement within 60 days.
Case support
Complaint exhibits for your works: catalog entries, MD5s, torrents, piece data and ownership screens, with a method note an expert can check.
Expert support
Your expert gets our scripts and checksums to re-run the method from the public sources. If you need a declaration on how the data was built, that is available as separate expert work, billed hourly.
Defendant lists
Lists of works by defendant and theory, as spreadsheets, for intake and case planning.
Client code packs
Codes your clients redeem for reports marked "Prepared for" your firm.
Who we work with. PayAuthors is independent research. We work with counsel on either side of these cases, and with authors, publishers and other rights holders, using the same data, method and standards for everyone. Custom work for a firm is confidential to that firm. Firm work is done under a short written engagement.
Five free reports
Request a code for five free reports.
Pricing
The name check is free. Author reports are $65 (Standard) and $100 (Registration Details); a Standard report can be upgraded for $35. The firm pilot is $1,000 for up to 50 works. Case support, defendant lists, code packs and expert declarations are priced per engagement.
Method and limits
We never download, open or share a pirated book, and never join a BitTorrent swarm. Every source file is recorded by SHA-256 checksum, and each registration match links to the Copyright Office record so it can be confirmed. Matching is automated and not perfect: expect some missed records and some matches that are not your clients'. A catalog listing shows that a file was in a library on a date; tying it to a particular download takes the defendant's own evidence. PayAuthors is not a law firm, does not refer clients to lawyers, and is not affiliated with or endorsed by the U.S. Copyright Office. See the Methodology for sources and checksums.
This page is general information, not legal advice.