PayAuthorscom

Methodology

Methodology and sources

For lawyers, experts and journalists. Dataset build payauthors-2026-10-01-m63 (v6.3 matches). Court records last reviewed October 1, 2026. The binding statement of limits is in the terms.

At a glance

dated catalog or listperiod of records with dates addedregistrations covered

How a registration is matched

  1. ISBN on both sides, confirmed by a surname, pen name, publisher or most of the title
  2. Exact full title + an author or claimant surname
  3. Exact main title (before a colon, slash or dash) + surname
  4. Same title words, stop words ignored, + surname, only when nothing else matched
  5. Exact full title + publisher named on a company registration

Strongest first. Each match in a report names its tier. Full rule

Measured accuracy

94.0% / 96.0%

Bartz opt-out works (Ex. J) found in the catalogs1,639 of the list’s 1,749 TX and TXu registrations; linked to the exact registration (1,540) / to some registration of the work · v6.3 rule

93.1% / 96.9%

Publishers’ listed textbook registrationsthose found in the catalogs, of 7,001 listed (mainly Cengage v. Google, ECF 1236-1): exact / some registration of the work · v6.3 rule

91% / 72%

How often a returned match is rightrandom sample of 100 matches (n = 100), checked against the Office record in a software review: same work / exact edition, weighted by method (about ±6 points)

How this was measured

Sources

Full description of each source
  1. Library Genesis catalogs, 25 snapshots, 2015–2025. The non-fiction ("updated" table) and fiction database dumps that Library Genesis published, as preserved by the Internet Archive on the dates below. Each catalog entry describes one file; entries describing the same edition are combined into one line, and each line carries the list of snapshots in which it appears, with the count of entries in each. The April 2019 snapshot contains the non-fiction catalog only. The February 2020 snapshot combines the non-fiction dump of February 13, 2020 with the fiction dump of February 8, 2020, which survives only as a Wayback Machine capture of libgen.rs/dbdumps/; both are dated February 13, 2020 in reports. Only catalog metadata is used; no book files are held.
  2. Pirate Library Mirror (PiLiMi) index of Z-Library, 2022. The metadata table of the PiLiMi mirror of Z-Library (11,783,153 records), in the metadata-only copy re-hosted on Hugging Face as P1ayer-1/annas-archive-index. Each record names the PiLiMi release that carried the file: 2,179,884 in the first release (July 2022, the release the Bartz v. Anthropic order describes Anthropic downloading that month), 3,818,910 in the second (September 2022), and 5,784,359 not mirrored (most already in LibGen). This index carries no ISBNs, so PiLiMi records match registrations by title and name only. No files from the mirror are held.
  3. Z-Library records added after PiLiMi, 2022–2025. Two later catalogs of Z-Library. (a) Anna’s Archive’s index of Z-Library records (“zlib3”), in the metadata-only copy on Hugging Face (P1ayer-1/annas-zlib3-index): 2,630,955 records with Z-Library IDs 22,430,000 to 25,639,999, added to Z-Library from August 24, 2022 to August 6, 2023, each with the date added, title, author, publisher, ISBNs and the MD5 Z-Library reported for the file. (b) The Z-Library extract in the MajinBook catalogue (Mazières and Poibeau, Zenodo, CC-BY 4.0; zlibrary.jsonl.gz, 8,097,488 records), which its authors describe as Z-Library’s metadata as of January 2025, obtained through Anna’s Archive, limited to EPUB-type records (EPUB, MOBI, AZW and similar); it carries IDs, titles, authors, years, languages, publishers and ISBNs, but no dates or MD5s. Records already in PiLiMi are not repeated (a PiLiMi line notes whether the January 2025 extract still lists it); 2,352,562 records are known only from the extract, and their IDs place them after August 2023. Anna’s Archive is a search engine and mirror that combines other collections; the court record on Meta’s early-2024 download of it is summarized below. No Z-Library or Anna’s Archive files are held.
  4. Internet Archive lending library. The catalog of the Internet Archive’s lending library (its inlibrary collection of scanned print books that signed-in users can borrow), read through the Archive’s public scrape API on the date given under Provenance: identifier, title, creator, ISBNs, publication year, publisher, language and the date the Archive added the item. The Internet Archive is a nonprofit library. In Hachette Book Group, Inc. v. Internet Archive, 115 F.4th 163 (2d Cir. 2024), the court held that the Archive’s scanning and lending of the 127 works at issue was not fair use. That AI developers obtained Internet Archive books is alleged in pending cases against Meta and NVIDIA; no court has found it.
  5. Journal articles (DOI lists). For the free article check: the list of 62,835,101 DOIs Sci-Hub posted on March 19, 2017 (Internet Archive item scihub_doi_list_2017, doi.7z); the DOI column of LibGen’s scimag database as published in February 2020 (libgen-meta-20200213, scimag_id_doi1_doi2_isbn_md5_pmid_pmcid.tsv.gz); and the scimag database dump of August 6, 2021 (libgen-meta-20210806, libgen_scimag_dbbackup-2021-08-06.rar), which also gives the date each article was added. Each DOI is lowercased, stripped of any doi: or doi.org/ prefix, and stored only as the first 8 bytes of its SHA-256 hash with a source mask; the article files were never fetched. Article lists for a person come from Crossref’s open metadata API (by ORCID iD where the publisher deposited it, or by author name). The free check shows list membership only. The paid Journal Article Registration Report (v6.2) adds two kinds of U.S. registration that may cover an article. Publisher side: the journal issue registered as a serial, from the Office’s serial registrations (the tabular file reg_serial_2026_01.csv, registrations 2006–2025, and the serial records of its MARC XML bulk file, registrations from 1978; 3,504,756 issue registrations across 174,707 serials). An article is matched to a registration through the journal’s ISSN (or its exact title when Crossref gives no ISSN) and the issue’s volume and number, or its volume or month where the registration gives only those; a serial registration covers the articles first published in that issue that the claimant fully owned when it filed (Circular 62B). Author side: a TX registration whose main title equals the article’s, with one of the article’s author surnames among the registration’s authors or claimants and a publication year within a year before or three years after the article’s; registrations carrying an ISBN (books) are left out. The § 412 timing condition is then compared with the dated article-copying events listed on the report, each graded. Many authors assigned their copyright to the publisher, so the publisher may be the owner who can sue.
  6. Books3 file-name list. A list of 196,608 Books3 file names published by Peter Schoppert (psmedia/Books3Info, filenames.txt). Books3 as released is reported to contain 196,640 books; the published list is 32 names shorter. No Books3 texts are held.
  7. Books3 ISBN list. 129,635 unique ISBN-13s that Schoppert extracted from the Books3 texts (Books2_ISBNs.txt, same repository). His README describes a pattern search over the texts that collected ISBN-13s only.
  8. U.S. Copyright Office registrations, 1978–2025. The Office's public tabular data for non-dramatic literary works (TX) as published in January 2026 (reg_non_dramatic_literary_work_2026_01.csv, 9.28 million rows, registrations from January 1, 1978 through 2025). These are public U.S. government records; PayAuthors is not affiliated with, or endorsed by, the U.S. Copyright Office.
LibGen snapshot dates and Internet Archive items (table)

LibGen snapshot dates and Internet Archive items

November 14, 2015libgen-meta-20151114
December 26, 2015libgen-meta-20151226
January 9, 2016libgen-meta-20160109
June 17, 2016libgen-meta-2016-06-17
November 4, 2016libgen-meta-20161104
December 30, 2016libgen-meta-20161230
May 12, 2017libgen-meta-20170512
April 15, 2019libgen-meta-20190415
February 13, 2020libgen-meta-20200213 (non-fiction); fiction dump of February 8, 2020: Wayback Machine capture, February 17, 2021
February 12, 2021libgen-meta-2021-02-12
June 5, 2021libgen.rs-dbdumps-libgen_2021-06-05.rar-2021-06-05 and libgen.rs-dbdumps-fiction_2021-06-05.rar-2021-06-05
March 7, 2022libgen-meta-20220307
March 8, 2022libgen-meta-20220308
June 19, 2022libgen-meta-20220619
January 13, 2023libgen-meta-20230113
April 29, 2023libgen-meta-20230429
May 13, 2023libgen-meta-20230513
January 6, 2024libgen-meta-20240106
June 28, 2024libgen-meta-20240628
October 2, 2024libgen-meta-20241002
November 16, 2024libgen-meta-20241116
December 17, 2024libgen-meta-20241217
January 18, 2025libgen-meta-20250118
March 8, 2025libgen-meta-20250308
June 3, 2025libgen-20250603

Each report lists the SHA-256 checksum of every source file it was built from, and the dataset build it came from. The full provenance table (archive item, file name, size, the archive's own SHA-1, our SHA-256, and the checksum of each parsed catalog) is published with each dataset build at the end of this page.

A checksum identifies the exact bytes of a file. By itself it does not establish a file's origin, completeness or authenticity.

Accuracy

Matching is automated and not perfect. Expect some missed records (other spellings, pen names, entries with no author, registrations under a different form of the name) and some matches that belong to other people with similar names or to other editions. Every line should be checked against the author’s own records and, for registrations, the linked Copyright Office record.

The searched name is split into parts. A record matches when every part is found at the start of a word within one author name of the record, so “Ann” matches Ann and Anna but not Joanne, and initials are optional. In the default Focused match the parts must fall within a short window of the same name, which excludes lists of other people’s names in the same field. Broader match drops that window and finds the parts anywhere in the author field, so it returns more records, including other people’s. Name order does not matter. Other spellings and pen names are separate searches.

LibGen catalog entries describing the same edition (same author field, title, ISBN, year, publisher and edition statement) are combined into one line, with a count of the entries; none are dropped. Books3 author and title are read out of the file name (“Title - Author.epub.txt”); where a name could not be split with confidence, both halves are shown. Some records cannot be found by name at all: about 138,000 non-fiction and 40,000 fiction LibGen entries carry no author, and about 19,000 Books3 file names carry no recognizable one.

File IDs and MD5 checksums

Each LibGen catalog entry describes one file and carries a numeric ID and the MD5 checksum of that file. A report line that combines several entries lists each entry’s ID, file type and MD5, with the snapshots that list the entry where they differ from the line’s. LibGen numbers non-fiction and fiction separately. PiLiMi records show the Z-Library ID and MD5 from the PiLiMi index. The Books3 file list has names only, so Books3 lines have neither. All of these come from catalog metadata; we have not seen or checked the files themselves.

More on file IDs, torrents, pieces, EPUB/MOBI and added dates

An MD5 identifies exact file contents, so an entry can be compared with a defendant’s file lists, hashes or logs produced in discovery. A match there would tie a work to a specific file. A catalog listing alone does not show that anyone obtained that file.

Torrents. LibGen distributes its files in torrents of 1,000 consecutive IDs, each named for the first ID in its block: non-fiction r_000, r_1000, r_2000 and so on, fiction f_0, f_1000 and so on. We checked this against the Internet Archive’s crawls of both torrent directories (November 20, 2020: 2,817 non-fiction and 2,223 fiction torrent names, every one a multiple of 1,000) and against a court filing that places two fiction files, identified by MD5, in fiction torrent “1000.” Our catalog lists both, as fiction IDs 1,286 and 1,291 (class plaintiffs’ Rule 56.1 statement ¶¶249, 251, In re OpenAI Copyright Litigation, S.D.N.Y. 1:25-md-03143, Dkt. 1987; an allegation). Each LibGen file line shows its torrent, worked out from its ID; confirm it against the torrent itself. PiLiMi lines show the PiLiMi torrent named in the PiLiMi index.

Torrent pieces. A torrent file lists every file it carries, with its size, and divides the combined data into pieces of a fixed length; BitTorrent peers exchange whole pieces. From the torrent files’ own metadata (file names, sizes and piece length; no book text), each file line shows the file’s size, the piece length, how many pieces the file spans, and the largest share of the file that lies inside one piece. A file that lies inside one piece is sent whole whenever that piece is sent. LibGen torrents come from the Internet Archive’s November 20, 2020 crawls and from 1,920 later block torrents captured one by one by the Wayback Machine at libgen.rs and libgen.is; PiLiMi torrents come from the PiLiMi torrent archive (pilimi-zlib-meta). Each LibGen file is located in its torrent by MD5 and each PiLiMi file by Z-Library ID. Piece data covers 99.6% of the non-fiction and 98.9% of the fiction files in the June 5, 2021 catalogs, and 95.2% and 84.0% of those in the June 3, 2025 catalogs, because some later torrents were never captured. Every PiLiMi release 1 file has piece data; release 2 was distributed as whole archive files, so its pieces do not follow book files and none is shown. Of the files located, 79% of fiction files, 26% of non-fiction files and 74% of PiLiMi release 1 files lie inside a single piece, and 97%, 61% and 93% have at least half of their data inside one piece. No torrent was joined.

EPUB and MOBI. File lines mark EPUB and MOBI files, because OpenAI’s 2018 download script is alleged to have taken only those two types (same filing, ¶147). The mark is information only; it does not change the statutory-damages timing mark.

Added dates. Each LibGen entry carries the date LibGen says the file was added, shown as “added to the catalog … (LibGen’s own record).” Between May 12, 2017 and February 12, 2021 the only fiction catalog that survives is the February 8, 2020 dump (a Wayback Machine capture of libgen.rs/dbdumps/fiction_2020-02-08.rar, 801,097,717 bytes, whose SHA-1 matches the digest the Wayback Machine recorded for the capture). Directory listings captured in 2019 show weekly fiction backups, but none of those files was captured. For fiction entries between those dates, the added date and the February 2020 dump are the only records of when an entry appeared. Entries deleted before a catalog we hold do not appear at all. As a check on the added dates: LibGen assigns IDs in order, and in our February 12, 2021 fiction catalog the added dates pass October 31, 2019 at about ID 2,190,000, below which the catalog holds 2,189,999 entries; the class plaintiffs’ filing puts LibGen fiction at 2,182,642 files in October 2019 (¶¶137, 236). In the February 13, 2020 non-fiction catalog the added dates pass January 31, 2020 at about ID 2,470,000, with 2,469,999 entries below it; the filing puts non-fiction at about 2.46 million as of January 2020 (¶137). Both agree within about half a percent (the ID cut-off is estimated to the nearest 10,000). The February 8, 2020 fiction dump allows a direct count: 2,185,033 of its 2,206,073 entries carry an added date on or before October 31, 2019, 0.1% above the filing’s 2,182,642, and 2,205,667 carry one on or before January 31, 2020.

How registrations are matched

Five tiers, strongest first, summarized under At a glance. Tiers 2 to 5 are heuristic: a match means a registration exists for that title and name, not that the catalog copy is that exact edition. Every match is a candidate to confirm against the linked Office record.

The full rule, the § 412 timing mark and its limits

Records are matched to Copyright Office TX registrations by the v6.3 rule (October 2, 2026), strongest first: (1) the same ISBN on both sides, kept only when the registration also shares an author or claimant surname, pen name or publisher name, or most of the title, because one ISBN is often printed on several editions, formats or an anthology; (2) the same full title plus an author or claimant surname; (3) the same main title (before a colon, slash or dash), or the book’s own title after a series name or number, plus surname, with safeguards against series names and disagreeing subtitles; (4) the same title words with stop words ignored, plus surname, only when nothing else matched; (5) the same full title plus the catalog’s publisher named on a company registration (a work made for hire). Titles are compared after routine normalization (case, punctuation, leading articles and edition notes), and common first names and role words are not treated as surnames. A match is set aside when the registration is for a translation, a study guide or an adaptation of the work (a play or graphic novel, say), or when volume, issue or edition numbers disagree. A record lists up to eight registrations of the work, chosen by how close their dates are to the record’s. Measured accuracy is below. Tiers 2 to 5 are heuristic: a match means a registration exists for that title and name, not that the catalog copy is that exact edition. Every match is a candidate to confirm against the linked Office record. TX registrations from 1978 through 2025 are covered, as published by the Office in January 2026; sound recordings such as audiobooks and renewals are not. Statutory-damages timing mark. For each matched registration the report compares the registration's effective date with each documented download date on the timeline (Anthropic: Books3 early 2021, LibGen June 2021, PiLiMi July 2022; Meta: October 2022 and after) and with the publication date in the registration. The screen sorts each comparison into one of four cases. (1) Known start: a court record, company statement or complaint gives when the download period began. (2) Catalog lower bound: the library’s own record of when that copy was added (LibGen’s or Z-Library’s record date; for Books3, its October 2020 release). This is catalog metadata as the library reported it, not independent proof of when anyone acquired the file. The earliest date the copying could have begun is the later of (1) and (2); for an event known only by an outer bound (a paper, model release or complaint date), only (2) is available. A registration dated before that earliest date passes the screen. (3) Publication route: a registration made within three months of a publication that came before that earliest date also passes (§ 412(2)). (4) Insufficient dates: a registration dated after the earliest date but before the end of the download period, or a comparison with no earliest date at all, is shown as “dates overlap or download date unknown” and never counted. The screen is deliberately conservative: it may leave out works for which § 412 would still allow statutory damages, for example a registration made within three months of publication where no earliest copying date is known. The condition comes from 17 U.S.C. § 412, which has exceptions this screen does not assess. The journal-article report uses the same four cases, with LibGen scimag’s record date as the catalog lower bound. For counts and dollar figures, copies, formats, editions and translations of one book count as one work, timed from its earliest copy in the dataset. It is a date screen, not a finding that the registration covers the listed edition, that the work was in the download, or that damages are available. Nor does it address whether a claim is timely: copyright claims must generally be brought within three years after they accrue (17 U.S.C. § 507(b)). In suits by authors who opted out of the Bartz v. Anthropic settlement, Anthropic argues (motion to dismiss, September 24, 2026) that claims based on its 2021–2022 downloads are time-barred and that the class action did not suspend the period for all of them; the court has not ruled. Limitations, the discovery rule, class-action tolling and the Bartz release should be checked with counsel for each work. Plaintiffs have begun to define classes by the same timing: the textbook-author classes proposed in Sullivan v. Meta (N.D. Cal., filed July 2, 2026) and Sullivan v. OpenAI Foundation (S.D.N.Y., filed August 14, 2026) cover only works registered within five years of first publication and before the defendant's copying, or within three months of first publication. Neither class has been certified.

Record-match marks

Every registration match in a report carries one of three marks.

In a review of 800 matches against the Copyright Office record, every match that passed the checks was the registered work: 272 the same edition and 5 another edition of it. The 800 were 500 matches drawn 100 from each matching method plus 300 drawn from random parts of the dataset; each was compared with its Office record in a software review that was not blinded to the matching method. The checks are strict, so many correct matches start as Not yet checked until reviewed. A record match says the registration is for this book; it is not a finding of ownership, copying or eligibility for damages. The report summary gives the count for each report, and the CSV has a registration_check column.

Measured accuracy

We test the matcher against works lists filed in court, which pair a title with its registration number. Of the Bartz v. Anthropic opt-out works (Dkt. 619-3, Ex. J; 1,749 TX and TXu registrations) found in these catalogs, the v6.3 rule links 94.0% to the exact registration in the court record and 96.0% to some registration of the work (v6.2: 91.3% and 94.8%). About four in ten of the remaining misses are entries where Ex. J itself pairs a title with another work’s registration. Of 7,001 textbook registrations in publishers’ works lists (mainly Cengage Learning v. Google, No. 1:24-cv-04274 (S.D.N.Y.), ECF 1236-1), it links 93.1% of those found in the catalogs to the exact registration and 96.9% to some registration of the work (v6.2: 90.7% and 96.7%). How often a returned match is right: in a fresh sample of 100 matches drawn from random parts of the dataset and checked against the Office record in a software review, 91% were the same work and 72% the exact edition, weighted by how often each matching method is used; 9% were another book (an ISBN printed on two books, a different book with the same title, or a translation that slipped through). The sample is small, so read these as approximate, about ±6 points. The 100 reviewed matches are published as a CSV (20 for each matching method, so the totals above weight them by how often each method is used: title and author 61%, main title 20%, ISBN 16%, loose title 2%, publisher 1%), each with its verdict, a note on any mismatch and a link to the Office record, so anyone can recheck them. The review is ours, not an independent audit. Earlier, on a random sample of 3,193 catalog records under the v6.2 rule, 99.8% of record–registration pairs passed an automated check that the title and an author, pen name or publisher agree; that is a consistency check between our own rules, not a measure of correctness. A review of 800 matches against the Copyright Office record is described under Verified marks. These are measurements on these lists, not a guarantee for any one match: confirm each registration against the linked Office record.

Copyright Office records, recorded documents and contact details

Each registration links to the Office’s own record in its Public Records System. The Registration Details report also lists every recorded document that cites the registration: transfers, assignments, licenses, security interests, notices of termination and others, each with its document number, date recorded, parties and, where the Office has them, addresses. These come from the Office’s bulk recordation data (published January 2026, documents recorded through June 26, 2025); later recordings are only in the Public Records System. The address on a registration is the one given at filing, so a recorded transfer can show that the claimant named there is no longer the owner.

For registrations from about 2010 on, the Registration Details PDF also reproduces the Office’s PDF of the registration record, unchanged, in an appendix. That record includes the claimant’s address and any rights-and-permissions contact the applicant supplied. Readers can use it to check whether those details look current. Supplementary registrations, which correct or add to an earlier registration, appear on the Office’s record but are not yet listed separately in our reports.

The “ISBN in Books3” flag

A LibGen edition is flagged when its ISBN appears in Schoppert’s ISBN list. The list does not say which Books3 file each ISBN came from, and a book’s text can print the ISBNs of other books or editions (for example in an “also by” list). The flag therefore supports, but does not by itself show, that the edition is in Books3.

We tested how often the two published lists agree. Of the ISBNs in the list, 98,035 could be identified through LibGen records. Of those, 66,661 (68%) matched a Books3 file name by both title and author surname; 8,310 matched the title but not the author; and 23,064 matched neither, more than half of them by authors who have other books in Books3. Where the ISBN and the file name agree, two independent derivations corroborate each other.

Snapshot dates and The Atlantic’s search tools

Each snapshot answers a timeline question: was an entry in the catalog by that date? With 25 LibGen snapshots from 2015 to 2025 a report can show when an edition was first listed and whether it was listed on the day of each documented download. The June 5, 2021 snapshot is the one closest to Anthropic's LibGen download; the June 2022 PiLiMi index precedes its PiLiMi download by weeks; the June 2022 and January 2023 LibGen snapshots bracket Meta's first use; the January 18, 2025 snapshot is from the same month as The Atlantic's. According to the Bartz v. Anthropic order, Anthropic downloaded at least five million books from LibGen in June 2021. The Kadrey v. Meta court found that Meta “first used a shadow library in October 2022,” about sixteen months after the snapshot. A match in the snapshot can help document that the catalog listed an entry before that acquisition. It does not show that a later defendant obtained that exact entry or file.

More on period events and The Atlantic’s search tools

For an event dated as a period (for example, OpenAI’s alleged torrenting from September 2019 through January 2020), the timeline also reports the first LibGen dump after the period ends, if it falls within 45 days, and how many of the lines it lists had a catalog entry added on or before the period’s last day (LibGen’s own record). For OpenAI’s alleged period that is the February 2020 pair of dumps (fiction February 8, non-fiction February 13), 8 and 13 days after the period ends. For the event dated only to 2018, the timeline also counts the EPUB and MOBI files listed in the snapshot compared, because the 2018 download script is alleged to have taken only those two types.

The Atlantic’s LibGen search uses a snapshot “taken in January 2025, after Meta is known to have accessed the database, so some titles here would not have been available to download,” and notes that “it’s impossible to know exactly which parts of LibGen Meta used to train its AI, and which parts it might have decided to exclude.” The Atlantic (a subscription may be required); the same caveat is quoted on Bruce Schneier’s blog. For Books3, The Atlantic’s Alex Reisner explained that the books are stored “as large, unlabeled blocks of text,” so he “extracted ISBNs from these blocks of text and looked them up in a book database”; of the 191,000 titles he identified, 183,000 had associated author information. The Atlantic, September 25, 2023.

The two approaches answer different questions. Ours documents what each dated catalog snapshot and a published file list contained, so a title added after a download shows up with the date it first appeared; a single later snapshot cannot separate those. A defendant’s own acquisition records, file manifests or logs would be more direct evidence of what that defendant obtained. Counts also differ because sources count different things: files, catalog records, identified titles, editions, or books with author metadata.

The court record

There is no single public ledger of every entity that copied every collection. The court records below are summarized as of the review date above; each is a statement about that case and its procedural stage.

The record, company by company

Huckabee v. Bloomberg (S.D.N.Y., November 24, 2025). The court denied Bloomberg’s motion to dismiss. The plaintiffs identified their works, alleged the works were in Books3, and relied on Bloomberg’s own paper saying it trained BloombergGPT on Books3. The court rejected the argument that the plaintiffs had to specify which files were theirs. This was a pleading-stage ruling, not a finding of liability. Decision.

Bartz v. Anthropic (N.D. Cal.). A June 2025 order described evidence that an Anthropic cofounder downloaded Books3 in early 2021, that Anthropic downloaded at least five million books from LibGen in June 2021, and at least two million from the Pirate Library Mirror in July 2022. The court treated building a central library from pirated books differently from using copies to train models. The case settled, with final approval in July 2026; the settlement does not resolve other companies’ conduct. 2025 order; final approval order.

Kadrey v. Meta (N.D. Cal.). The June 2025 opinion describes Meta downloading LibGen in October 2022 and using Books3 in training data, and granted Meta summary judgment on the training claims on that record; the judge said the ruling “stands only for the proposition that these plaintiffs made the wrong arguments and failed to develop a record.” Distribution through BitTorrent was left for separate proof. The opinion also describes Meta downloading Anna’s Archive, a compilation that includes LibGen and Z-Library, in early 2024, and says it is undisputed that Meta torrented LibGen and Anna’s Archive (Dkt. 598 at 11–13). In March 2026 the court granted leave to add a contributory infringement claim based on BitTorrent uploading. In its July 13, 2026 Answer in Elsevier v. Meta (S.D.N.Y. 1:26-cv-03689, Dkt. 78), Meta admits obtaining data from LibGen, Anna’s Archive and Sci-Hub to train some Llama models, and obtaining Anna’s Archive data in 2024 (¶¶82, 97); that is a company statement in a pleading, not a court finding. 2025 opinion; March 2026 order.

NVIDIA. A court allowed claims to proceed past dismissal on allegations connecting Books3, The Pile and a model-training dataset; that was not a finding that NVIDIA copied any particular book. The amended complaint alleges that NVIDIA sought high-speed access to Anna’s Archive in August 2023 and that management approved within a week, and that the offer included millions of Internet Archive lending books (ECF 235 ¶¶53–57). NVIDIA withdrew its motion to dismiss as to Anna’s Archive, Z-Library, LibGen and Sci-Hub (ECF 271, March 6, 2026), and its June 2026 Answer admits downloading data for AI research from annas-archive.org/datasets and libgen.rs, without dates, while denying that it improperly copied the plaintiffs’ works (ECF 321 ¶¶58, 77). Order.

DeepSeek. DeepSeek’s own DeepSeek-VL paper (arXiv 2403.05525, March 8, 2024, §2.1) says it cleaned 860,000 English and 180,000 Chinese e-books from Anna’s Archive for its document-OCR training data. That is a company statement; we know of no U.S. court ruling on it. Paper.

OpenAI. A November 2025 discovery order described as undisputed that an OpenAI employee downloaded books from LibGen in 2018; other details about derived datasets came from testimony and party assertions. Order.

Where a company’s own statements or a court record show it copied an entire collection, what an author then needs is the link between the work and that collection: the work, its registration, and the specific entry or file name in a dated source. That is the link a PayAuthors report is built to document. What the company copied, and when, has to come from the company’s own records, statements or court findings.

Provenance: dataset build payauthors-2026-10-01-m63

Checksums of every source file

Added in this build (v6)

Z-Library records after PiLiMi, the Internet Archive lending catalog and the journal-article DOI lists. The two scimag files were read as a stream and not stored; their checksums are the Internet Archive’s.

SourceFileSHA-256
Hugging Face P1ayer-1/annas-zlib3-indextrain-00000-of-00005-421d14d609b06e30.parquetc06f449014008b8a340e2867d410cca603600bed7b0789faf086e4d5f91ec062
Hugging Face P1ayer-1/annas-zlib3-indextrain-00001-of-00005-33b56dc710868fcc.parquet0a10340de0523a670668108181b37c1bbd6fdfe3d6b72e0ca778b30f0cdb885a
Hugging Face P1ayer-1/annas-zlib3-indextrain-00002-of-00005-8801c913706c9de7.parquet03f8fb4e2130d081025c1b0678a43e16fd496f590f336481694ccfe3ccd16e1b
Hugging Face P1ayer-1/annas-zlib3-indextrain-00003-of-00005-11a7abbf65bdb652.parquet6eddb91d0a0133654434d69824e74344dabdba8de43cdc459744e1612a8e8735
Hugging Face P1ayer-1/annas-zlib3-indextrain-00004-of-00005-9f57636ca249d4d4.parquetd4ee2a61f1d75b2ebd758638f6162089b4cabf8d258be21b92df22fd3345ef7e
Zenodo 10.5281/zenodo.17609567 (MD5 5876dc0cb3ef4be79a35fbaf59d5bcbe)majinbook_zlibrary.jsonl.gzda1d50b04c97f79c6c8ac1030f74cd0f083104e62da19a3bffe1e1ed626ad071
Internet Archive scrape API, collection:inlibrary, fetched Sept. 30 – Oct. 1, 2026 (4,423,866 items)inlibrary.ndjson.gz (our download)30d7c22621e6992cd323a5801617cd99762a309e1066a90c9228f8b99c2b0934
Internet Archive item scihub_doi_list_2017 (MD5 62b253ebdb9456a7e632e81d7e6ed800)doi.7z(MD5 as published by the Archive)
Internet Archive item libgen-meta-20200213 (MD5 bdcdb69511fe0aa328b1afc659ab5714)scimag_id_doi1_doi2_isbn_md5_pmid_pmcid.tsv.gz (read as a stream)(MD5 as published by the Archive)
Internet Archive item libgen-meta-20210806 (MD5 084992920a17ce3dc9fe65b69803339c)libgen_scimag_dbbackup-2021-08-06.rar (read as a stream)(MD5 as published by the Archive)
U.S. Copyright Office, data.copyright.gov Registrations/Tabular/ (v6.2, journal-article registrations)reg_serial_2026_01.csv (401,155,206 bytes)02e034580e8d638d7a3c8fc471861452d22ceaec725d7b9cb56b686effffc71a
U.S. Copyright Office, data.copyright.gov Raw unparsed uncategorized XML/ (v6.2; serial records only, read as a stream)USCO public data.zip (3,581,316,904 bytes, dated Apr. 2, 2024)(read as a stream; not stored)

Every source file behind the current dataset build. LibGen items link to their Internet Archive pages; “SHA-1 (archive)” is the checksum the Internet Archive publishes for the file (for the Wayback Machine capture, the capture’s recorded payload digest), which we matched before use, and SHA-256 is our own. Only catalog metadata was read from these files. The file IDs and MD5 checksums shown on report lines were collected in a second pass over the same files, each downloaded again and matched to the same SHA-1 and SHA-256 (the February 8, 2020 fiction dump, added later, in a single pass).

SnapshotCatalogArchive item and fileBytesSHA-1 (archive)SHA-256
November 14, 2015fictionlibgen-meta-20151114
20151114-fiction.rar
169,417,419b66dc4473c41935a5061174c6adca77815b79c39fb28917f8f8102782092ba37c53ed89823202d0b900530706a86cff7b2a3b6f7
November 14, 2015non-fictionlibgen-meta-20151114
20151114-libgen_compact.rar
335,440,447722d792dcef25dec518bba879be42990329d91410927f0659320624e0d00a8288cc54548bee676d244df7ae1055f507d33b9f506
December 26, 2015fictionlibgen-meta-20151226
20151226-fiction.rar
229,964,419b1575d11a43eddc8660a77383e53376a95b690ed2d8f895eca25dcc74e05221998b2cf06aa998dcb26c7ae4b7d37ca559ed850ea
December 26, 2015non-fictionlibgen-meta-20151226
20151226-libgen_compact.rar
343,348,0070e1a623afee6070b2c5566e25d3fb386f1f5d3b943a80461adb48fbf380bc3c385b0558c122d082cef02c1586fbf1d6c17452177
January 9, 2016fictionlibgen-meta-20160109
fiction.rar
230,176,42944a2dd92c6eedb1378613be529aca3bd23bea904c3aa13dfa7ddb4e842a3ba55dddbaf9fe30908831b5f4faba2e66631a3fb88f5
January 9, 2016non-fictionlibgen-meta-20160109
libgen_compact.rar
343,922,2236a627c137f3b54aa9b182e9d6b532905fc85a7a3598af05e5b821e7eaec45f8bdb8575639bd38eca5ed5620f0feb6d27a346edab
June 17, 2016fictionlibgen-meta-2016-06-17
fiction_dbbackup-2016-06-17.rar
240,314,699a9b07c67a247c69d90a56218ce01e3a5a9d941c03590b1ead76d2e0bd8247f917ccfe90d237744b05e62d085152f9544a062ba4c
June 17, 2016non-fictionlibgen-meta-2016-06-17
libgen_dbbackup-2016-06-17.rar
190,512,12215856da1e67fd0c630f6d39a8760dd6a9d50c30d3d54db847419b450df4e40830ffc95288230bf66c8c9ddc661e7809647dfb241
November 4, 2016fictionlibgen-meta-20161104
fiction_dbbackup-last.rar
241,894,3405bfb65ba9defb1c572c612d094e19b1fc126c7ec9db15b0c17e6f8e6cce4109512a187d50057c19ba525f27502b92408543a3d61
November 4, 2016non-fictionlibgen-meta-20161104
libgen_dbbackup-last.rar
199,327,8174609fce0f64bab73352b4e775e58de2747f920ef615b16cdcc31e849c2eab83b11bb54caa7343b710595c25dc9fefad5f7c62029
December 30, 2016fictionlibgen-meta-20161230
fiction_dbbackup-2016-12-30.rar
243,419,703f4a26cd213ffda944bea06e1e33c615646cdae97972a14c00e53de461f7ac9c3cbfc9a0739a5dd02ad3a4cf3588c2376fa40bb83
December 30, 2016non-fictionlibgen-meta-20161230
libgen_dbbackup-2016-12-30.rar
205,108,150acae4bc3f6952239fccc10f51a183e597f3044c518572f6615ffce8d5cb8503d52a4751ca3abb0b641572ee7e368599da43a8447
May 12, 2017fictionlibgen-meta-20170512
fiction_dbbackup-2017-05-12.rar
456,237,0438d9be107c93d3e4bf454f4c6cd8720210de01a1ffdfe5d94a98ffd1c39e091fb99815a6998a287e572e8bc74c581a08345f19ff2
May 12, 2017non-fictionlibgen-meta-20170512
libgen_dbbackup-2017-05-12.rar
213,571,994db4d349b5e22a672dac1828369353c3a683374c0e37193c21088da07c8e6a4f1290dc66678caf13975ce658c26bcf587148fa3aa
April 15, 2019non-fictionlibgen-meta-20190415
libgen_compact_2019-04-15.rar
282,762,38800f3772fc4942f7a48bc8c1b91a9de8880e6d877b5fae3c84a6a40512371dd6a8a26ab422feb67c6184283f6e64568f36e6f4660
February 8, 2020 (folded with the February 13, 2020 non-fiction dump)fictionWayback Machine capture of 2021-02-17
http://libgen.rs/dbdumps/fiction_2020-02-08.rar
801,097,717c0761e9e7d8fe98d32a98195f89e564a37cd41c6f28145748b8cb9f730eb879c406638ebe78f43423b7b99217b6e90800e6a6bd4
February 13, 2020non-fictionlibgen-meta-20200213
libgen_compact_2020-02-13.rar
269,446,563dbf10c896c4d30c5284f3ec6229990a0d7cdee309fb734297d87765fbf2332f139dda4cf56d2f2f2cf0c2ec91425f75380259363
February 12, 2021fictionlibgen-meta-2021-02-12
fiction_2021-02-12.rar
926,040,033fed262df72583676e838ae3a18ec9fc205065d73b2ff2775dff911f3e9f6599d18d34a4ee4a714ac4f9e76513e0e2c1013a15dec
February 12, 2021non-fictionlibgen-meta-2021-02-12
libgen_compact_2021-02-12.rar
328,162,4535b93fb99ff9b821387ff8f12c63dfebd651f0589bef9a8430de38848561849d7b93aeae06ee0c2b62b9658e3949da281871f60fd
June 5, 2021fictionlibgen.rs-dbdumps-fiction_2021-06-05.rar-2021-06-05
libgen.rs-dbdumps-fiction_2021-06-05.rar-2021-06-05-b4787ae8-00000.warc.gz (Internet Archive web capture)
934,951,854f1ed9bc84cc931cd72c9abe1397617de5670c56f8c9b782a51b4fb358ad3118c6eaf10a622576b55c21536c80ef3b0c13f59fd5b
June 5, 2021non-fictionlibgen.rs-dbdumps-libgen_2021-06-05.rar-2021-06-05
libgen.rs-dbdumps-libgen_2021-06-05.rar-2021-06-05-a77164ac-00000.warc.gz (Internet Archive web capture)
4,075,381,0251b459a44a83ee0ab0505c1c43a8da857d1c5c8b46f40db38486bad5250cbef257b74a23e5cd25837729be9713910b24a7d4020b2
March 7, 2022fictionlibgen-meta-20220307
fiction_2022-03-07.rar
985,753,7835e9b00e36162480f70fc8971bf5bd7ab79521de030ac56c795a57ace4bd9d549bccb61ce45bf31fc0860cb99cd6d79bf3d596d00
March 7, 2022non-fictionlibgen-meta-20220307
libgen_compact_2022-03-07.rar
367,472,6328e906f5f22ccdac807c23c015e6fbbebd6d0d186e2645975a21a0507365b886f1d7eb61f3d3f9690a4912150b0958319099de3ed
March 8, 2022fictionlibgen-meta-20220308
fiction_2022-03-08.rar
985,829,360f24216a4dca7c7d0fecdb5d424644f4d885d3dd0a68eee73aa4337f5a5126da32f9868b6a2700bf078c9c0fd00300d4c50f48318
March 8, 2022non-fictionlibgen-meta-20220308
libgen_compact_2022-03-08.rar
367,637,914bafbe9add97c7fdcffe9d64f27e1d8f4ce8c3b6c0cc175e41f042f31f485b2b669a52c258ddc2ed40e43b82ad479a87f437fde35
June 19, 2022fictionlibgen-meta-20220619
fiction_2022-06-19.rar
997,491,750a0b5ab9a0074f381f8fd729e59cb1fe55db1e583cdb8a821e702785da84ddb448fdf92bb291e06f30aeb3a20792e773d9b7bbc12
June 19, 2022non-fictionlibgen-meta-20220619
libgen_compact_2022-06-19.rar
376,799,809aa9a072af84d2315d0e2c3e0ca1eb119dee66617c989f6e8300a444c01b2bb99ac21d51a3aca2b3a635366cc6a42751310b9ef56
January 13, 2023fictionlibgen-meta-20230113
fiction_2023-01-13.rar
1,103,284,2839d1724ad11bf03053c7e5d7ee6fa105a7caa9b4b197b72839f0d9665726c8cbd4168482e1c42654926a55d6ed109f7d7541c8a3e
January 13, 2023non-fictionlibgen-meta-20230113
libgen_compact_2023-01-13.rar
415,009,403e597c13ff1455eb7efeeb5f29275c019aeb967505f87a00eefba1e93f858bf74c96a5a9a1d0f9aa5188415202e7c247b6921c367
April 29, 2023fictionlibgen-meta-20230429
fiction_2023-04-25.rar
1,133,284,237b6c71aa3940bd795c741124b957f57b89b3dc4bcc5aa780f89412ba8fbbfff8b75bc17a0276e249b27b75ff97d864860e912b508
April 29, 2023non-fictionlibgen-meta-20230429
libgen_compact_2023-04-29.rar
429,176,031f178bc51cfb942d56891275c2b0a549230b32a72b4dbb18e038b0f4229b4ab2b957993ab42bb116b48c56e8f5ba715af5083eee1
May 13, 2023fictionlibgen-meta-20230513
fiction_2023-05-13.rar
1,138,348,9603dc1076c8f7b5c231423576b39c134dcc6aa5cd1f2c488fd745fbd3ca3d2232ecb5b9c64c8f709dbe077f1e758d0d548cdc4e028
May 13, 2023non-fictionlibgen-meta-20230513
libgen_compact_2023-05-13.rar
430,711,9350442b5658304a837677a5f7b5f916d3fbe8d06a154705986f983d0dfcec53846113d1f2faa7fe46f0438331357c985dc83d15506
January 6, 2024fictionlibgen-meta-20240106
fiction_2024-01-06.rar
1,188,758,7693364b11ac9c02353d9c14162e6606252e9199dc7be3ee7bf05f824af431fbe821cd3783894438bace8f09ca5936e0d7d06caa0e7
January 6, 2024non-fictionlibgen-meta-20240106
libgen_compact_2024-01-06.rar
488,627,470a121566db551406e62acdcfc79ca339a8943da53fcbd99d7fa18e67ef7732ad03316e6a723ff8556149ac9789c87c160817e4bc5
June 28, 2024fictionlibgen-meta-20240628
fiction_2024-06-28.rar
1,261,925,704262fcf82d4e93d15d74c85860d4bfe35026786443456d0a3199df8735eb42f50210baa275ef0344fcff5d7a8bcb12cbc0f417b32
June 28, 2024non-fictionlibgen-meta-20240628
libgen_compact_2024-06-28.rar
507,660,3822daaa7353a3283ef859d0c0e8f7f8cc90fe1a3654e3746afd983e61521956e4b1c6783fba58072a2b3f01cc88bfeffb94ed6dece
October 2, 2024fictionlibgen-meta-20241002
fiction_2024-10-02.rar
1,278,582,785296dc599358becc9b0dac7b9f09c7998b967845e7cdd3d2b92cf80b86013f1d8cd3e24f3f689a3ab5bcfd0a627a20327784bc1c9
October 2, 2024non-fictionlibgen-meta-20241002
libgen_compact_2024-10-02.rar
514,088,078d1336ba78c25353bf69551c627684a7951625e17a8bd80490de629733e61419fbf5ad063ad3565da92cae621d2c6b187f8b982b8
November 16, 2024fictionlibgen-meta-20241116
fiction_2024-11-16.rar
1,288,550,033da2faa57b3f603b7125fde1dd5a4ff793b58262592ead19d68e4b992c3e9c935bb4d6a5cc653bed6eed44ed237d3e88143fce202
November 16, 2024non-fictionlibgen-meta-20241116
libgen_compact_2024-11-16.rar
518,055,73644bbbf2cde48ee61bd127e05a904f9b67b4232a68fd4ef4aa5c97326ddadb3dc70f11723d33a76f26c9b4b7aea93be3f357516a0
December 17, 2024fictionlibgen-meta-20241217
fiction_2024-12-17.rar
1,296,846,448557c037cb4b75d26b7e361ffba5faa7d92eb0fda8bba3c8f90a2800a32252253501a98d0c5a311479a7527ffd9ac41846e58695b
December 17, 2024non-fictionlibgen-meta-20241217
libgen_compact_2024-12-17.rar
520,829,3184a95bd454e455a3416b29a05c2ab26ca7020e74c87a74ca74c07f9766c9e5392eacf6a6d354b4edce9983225e4ba19fbd8dcb4a0
January 18, 2025fictionlibgen-meta-20250118
fiction_2025-01-18.rar
1,304,646,8785e51ccfba19e2f13bfdda653dae4429ef677cbd5f43a5f4e0e7782964d50c3e8b9f47d86abc3cd6039cfffb135e32ef14596c7c9
January 18, 2025non-fictionlibgen-meta-20250118
libgen_compact_2025-01-18.rar
522,699,977208a4fa0dfa4aa06308c253d968a7aec5c7066f91dd45de89caafacbe9f1b412d027fbf2807e955924e2799348906e9c4460a117
March 8, 2025fictionlibgen-meta-20250308
fiction_2025-03-08.rar
1,313,545,516cedd7c973133431b5b69dbf02bdf1cbee53c80dd946a7efa87fa51fd436b0117db51b290d16ca954558b10729679fb2f3ed725a4
March 8, 2025non-fictionlibgen-meta-20250308
libgen_compact_2025-03-08.rar
524,887,099f8ec8baaf8a4405c7543093535ef9dbe531cb9813cc299f25c24ee7ba75c9adc890aae5c9aed8c3ceebf8a71eb9c72a554306299
June 3, 2025fictionlibgen-20250603
fiction_2025-06-03.rar
1,324,697,8870b67a4117d62dcdd0b2e63691d6c6ba8f9315bf6e22c00ea1d31b7d154e76813552b72aba480461d919c4bf843465fbef6180e26
June 3, 2025non-fictionlibgen-20250603
libgen_compact_2025-06-03.rar
530,821,7637b763fdc80b1e6cc3f4c6def3bcb2fe9b8067024e3e702e6496364073cb38f536efc43b325606690f76cba4c18cf10904075f5b2
SourceFileSHA-256
PiLiMi (Z-Library) indexMetadata-only copy re-hosted on Hugging Face, P1ayer-1/annas-archive-index, 18 parquet files; records tagged by PiLiMi release (July and September 2022)b3005a74ca30e2ca6dc5a8ddc8a275db340f3877271672ae752555b9d8f08f15 train-00000-of-00018-0de08d15f83d5a95.parquet
086083645d2dee61e9ca8e069de6f5dfa854a96b63417a60e72a0f5d8e339907 train-00001-of-00018-c5b87dfcfaaca263.parquet
4d8dc8fdd98d0b065eed786c2756873b652a8383a62fbb76805f113276ecc7a0 train-00002-of-00018-b0dd6be7e63187a1.parquet
f731700cb6681f16b5d7011d33b6cfe29352daf6ac15f38c54662db0f9586f37 train-00003-of-00018-0bbb9ba4be152d8b.parquet
dc27ecd99d3245632291262d7957e579cff2a2929b8324891b3c5d2674b207d8 train-00004-of-00018-ea10c700bc0c6835.parquet
e44756886570f26764a317b66345cb85166a872022d87ba1aafe89750dc1d350 train-00005-of-00018-fa58b4f7bc3ec283.parquet
47505e79575f144f45eefeae30e882fcd08e0d58841aadf397911072c6cf8971 train-00006-of-00018-ed59a37788bfedc5.parquet
7f811cfd2ae5a2d03ada29221430af4875b1602c38bae298bbde1d0885e85370 train-00007-of-00018-e36786f5069ebd1d.parquet
2ae0a906abe3ec0538f70f7948556800ecb264160b9bfe338c592a3cd36a96e2 train-00008-of-00018-91a3071dc8471e38.parquet
1ee01c5c8864d43a1393eca3adc49fc16639b1ccadf88366c0cec1d035d2ac40 train-00009-of-00018-0b5b1757686f33bd.parquet
500159efccc2fe74ec365f244843c92d2eb9d81b68dc7dacab7c390e4363f2cf train-00010-of-00018-23cd38004aa9fe5e.parquet
b1046096da29e2aff76a45951bbc96bfeff2f1029aa425823e99e7f4a583927a train-00011-of-00018-f0423707a2f379dd.parquet
60b3d2c0ee333db27437ab87d8c8b6cb1b9a5dcd378e0d30d2cf546fe6084667 train-00012-of-00018-3541a214f73c36cf.parquet
e851a1ef375b8a71fd39563cd733b3fc619a3ddaf5e9589abe7a9e4cc3e3e24f train-00013-of-00018-94ea5b888df94a3e.parquet
5242e4917d79406d4a57dacb93874b8931051fe114901f55aee61cc67ca02c7f train-00014-of-00018-ff19e1d3c674dce3.parquet
048cdf4b90fcad0142968b56defe7ed43a26b27d2f828882173109da14912938 train-00015-of-00018-dd633849fdd8070c.parquet
90d4903cb35e4f9fd5ddb9542ed57d046c8c4e18df5e76af09c200862921fe83 train-00016-of-00018-cb1438e48baf670a.parquet
858089c6992df6962bfb05cef701558ba6fe4360dea77c0b63953ea7a7a79c5e train-00017-of-00018-b2b51d65a0cfde19.parquet
Books3 file listfilenames.txt, github.com/psmedia/Books3Info (Peter Schoppert)eb34489806bc36674a36515f9d5eb1d07a06edbd53ef2ce6aeb532e027bd7faa
Books3 ISBN listBooks2_ISBNs.txt, same repository (129,635 ISBN-13s extracted from the Books3 texts)6d67cb232fc60b10e34fe0afc7ffb409528f519b49d57cc79647179fd9ca99e3
U.S. Copyright Office registrationsreg_non_dramatic_literary_work_2026_01.csv, data.copyright.gov Registrations/Tabular (non-dramatic literary works, as published January 2026; 4,789,987,093 bytes)4bfba9f70bf47954d964f9350c495c4bc953149521a1e25b10f3bc9d2b925336
U.S. Copyright Office recorded documentsrecordation_2026_01.zip, data.copyright.gov Recordations (published January 2026; documents recorded through June 26, 2025; 1,630,294,321 bytes)99ec235192fd03203e7d6f95466774f3f5f07cef37ac7d606fd196425cd707a9

This page is general information, not legal advice. Court records and source datasets may change or be interpreted differently as litigation proceeds.