People say: if Pangram used the same book dump everyone else had, there would be a BitTorrent log on a Brooklyn office subnet and a journalist would have it by now.
That is not how this industry did the crime when the crime was large enough to leave mail.
In April 2023 a Meta research engineer, Nikolay Bashlykov, wrote the sentence that should be on the slide: “Torrenting from a corporate laptop doesn’t feel right.” He asked whether they could load LibGen on Meta IP ranges or should use a VPN. A later internal note, the Kadrey plaintiffs say, decided not to use Facebook infrastructure for the download, to “avoid risk of tracing back the seeder/downloader from FB servers.” They called it stealth mode. They turned seeding down to the smallest setting the protocol would allow. They still, the authors allege, took 81.7 terabytes through Anna’s Archive and had already taken 80.6 from LibGen. Meta has called the training fair use. The concealment is the tell. Nobody who intended to be proud of the catalog used the office IP.
Anthropic’s version is in a published judicial order. In June 2021 an employee downloaded at least five million books from LibGen, “which he knew had been pirated.” In July 2022 the company downloaded at least two million more from PiLiMi. Alsup: more than seven million copies into a permanent central library, not excused by later buying a bookstore copy, not excused by the training being transformative. That library was not a press release. It was a disk.
So when a detector startup in 2024 writes “licensed for commercial use” and “not unauthorized internet crawls,” and then prints Books: 7,000,000 examples without a bibliography, the naive read is Gutenberg. The adult read is: this is the sentence you write if the files arrived the way everyone else’s files arrived, and you are not about to put the filename on a Series A slide.
What we actually have on Pangram
Fact: the 2024 paper’s Table 5 lists seven million book examples in a 28 million document human pool, dated 2021 and earlier, “open source” and “licensed for commercial use.”
Fact: they evaluate the books domain on Project Gutenberg. Gutenberg’s whole catalog is on the order of seventy thousand titles. You can cut seventy thousand books into seven million chunks. You can also have seven million book-shaped rows from a shadow library. Those are different objects. They do not say which.
Fact: the training recipe is a human page versus an LLM “synthetic mirror” of the same topic and length, sometimes starting from the human’s first sentences. That recipe is cheapest if you already have a pile of books and a pile of API keys. It is the dump-shaped idea even if every file was a librarian’s gift.
Fact: Pangram 4’s card repeats the safe sentence: owned, commercially licensed, or openly licensed; no user submissions; no unauthorized crawls. That is the legal description of not-Books3. It is also what you say when you do not want a Kadrey-style discovery letter.
Not a fact: Pangram torrented Anna’s Archive. There is no unsealed Slack. A two-year-old startup does not leave 81 terabytes on a letterhead. A two-year-old startup leaves a hard drive somebody filled on a home connection, or buys a copy of a corpus that already made the rounds, or chunks Gutenberg and hopes you will not ask for ISBNs.
The assumption, stated as an assumption: they used the same book pile the labs used, or a descendant of it, and they will not say so because it is frowned upon. Frowned upon is doing a lot of work. It is frowned upon the way seeding from a Facebook server was frowned upon. Not “we would never.” “We would never get caught on the office IP.”
Why the assumption is the reasonable prior
- Every lab that mattered and got sued had the dump. Anthropic’s number of pirated copies — over seven million — is sitting next to Pangram’s seven million book examples like a rhyme. A rhyme is not a receipt. It is a smell.
- Hard-negative mining, their Algorithm 1, searches a “large” human reserve they do not list. The algorithm is: find the humans you already fail, generate mirrors, retrain. The reserve is the interesting object. They will not name it.
- Gutenberg, Books3, LibGen, and Anna’s Archive share a grotesque amount of the same dead authors. Overlap is free. If their human class loves 1925 sentences — and in our scans, leftover-free Gutenberg Keep is 14-for-14 entire-text Human — you do not need a new theory of stylometry. You need a theory of whose books sat in the train set.
- The people who hide downloads do not use corporate accounts. That is no longer a hunch. That is Meta mail.
If you work at a detector company in 2024 and you want a book-versus-rewrite classifier, you do not start by calling Penguin’s licensing desk. You start with what is on disk in the community. Then you write “commercially licensed.” Then you raise from the firm that already bet on the lab that torrented the books.
What would kill the assumption
A bibliography. Seven million book examples. Names, licenses, years, whether a row is a whole book or a 512-token slice. If it is Gutenberg plus a named licensed catalog, I will write the correction in the same type. If it is a private deal with a distributor, say so. Until then, “we would never” is not evidence. It is the press sentence of every company that later produced a torrent client log.
A living-prose control would also help. Score a thousand PERSUADE or ELLIPSE essays that are not in LibGen. If they come back Human, the detector can see a person who is still alive. If they come back AI, the “one in 10,000” story is a story about book-shaped English, and the dump question got less hypothetical.
I will not write “Pangram torrented 81 terabytes.” That number is Meta’s, in a complaint. I will write: the method is the dump, the catalog is unnamed, and the last time this industry had a catalog like that, they asked for a VPN. Part 3 is what that does to people who still type.
Sources: Ars Technica, 6 Feb 2025; TorrentFreak / Kadrey; Alsup order in Bartz (June 2025); arXiv 2402.14873 Table 5, Algorithm 1; Pangram 4 model card.