Aug 9, 2026
The Story
AI companies seem to have been pirating books, then ingesting them and training their models on them. Since 2024. Or maybe the whole time. Why exactly?
Since after all the public Internet has already been ingested by AI models, there is no new fresh quality - just AI-generated slop that gets looped back (e.g. lots of new content is AI-gen), leading to what people call model collapse LOL.
Last month Anthropic paid 1.5 billion dollar fine for doing the pirating, allegedly starting in 2024:
I doubt they are the only ones with this filthy practice.
However, while piracy is an outlaw activity, it appears buying books and doing the same (AI model training) is not against the law - not considered theft or IP infringement. Interesting...
It seems Dario is hungry for the world's knowledge.
How the scheme works? It is simple: you are an AI company, you go to somebody (anyone) else to buy some books for you en masse. Then nobody can tell it was you! Simple AF. It reminds me of the purpose of shell companies offshore, laundering money. But it's now laundering data... or pages.
My guess is if you are a normal person watching how AI corporations have been behaving over the last year and a half, you wouldn't be surprised at all. You would simply be disgusted ones again. They will simply do EVERYTHING they can for your data, for the world's data, just to lock it up behind a paywall. For the possibility to rent it out back to you. Ofcourse, for a small fee (until it isn't that small).
But wait, there is more. Dario is both hungry and doesn't want others to have it. Being like:
Ok, scan the goddamn paper and throw it in the trash
WAIT, WHAT?! Why throw it in the trash?? Dario continues:
Yes, that's right - the technology to scan faster includes the chopping of the spine of a book. We don't really care. We not only don't care, it's to our benefit! Let's make our models the ONLY ones that have those rare pieces of knowledge.. Mua.. muaa.. MUAAAAHAHAHAHAHAH (evilish laughter)
Chopping off the spine... oh boy.......
If you are a more of an analog person like me, you would be in pain just reading this.
Anthropic used a method called destructive scanning - buying physical books in bulk, chopping off their spines with hydraulic cutters, feeding the loose pages into high-speed digital scanners, and recycling the paper - to gather high-quality text data to train its Claude AI models.
If you want more details on the entire subject, check out this nice article by the Atlantic. I couldn't agree more with this one quote:
Even if a literal book burning is not under way, the process of ingesting millions of books, stripping them of authorship, and blending them into a homogenous “intelligence” branded with the name of a chatbot certainly feels destructive.
Poetry
The Paper Sacrifice
A thousand pages, neatly laid,
The spine will give way to a blade,
Wisdom reduced into digital stream,
A cold but efficient AI corp wet dream.
The Legal Loophole
Buy it. Scan it. Pulp it down.
Under US law, the crown is found.
First Sale Doctrine, a loophole so wide,
Where history's physical record must hide.
Haiku
Paper to the code,
Knowledge chopped, then digitized,
Wisdom is consumed.
The hungry machine,
Eats the rare, the niche, the old,
Nothing left to keep.
The Implications
So, what does this mean for the rest of us - analog, book-loving poor souls, who are not always satisfied with a regurgitated chat box response?
I don't see this as a quirky corporate data grab but a systemic shift in how the knowledge of the entire civilization is valued. The AI race has redefined "value" from cultural preservation to raw, clean data density. Data becomes more and more important - the ability to acquire and "own" it in non-polluted high-quality format is so important to those companies that they are willing to destroy it rather then preserve it (in line with imperialistic tendencies). Forget about ethical and cultural concerns provided the legal loopholes exist. The First-sale doctrine acts as a shield, allowing companies to treat cultural artifacts (such as books) not as heritage but as disposable one-time raw material.
While voices of European booksellers and preservationists - would be sounding the alarm, the existing legal framework is simply permissive enough. The destruction will keep happening, but whether any regulatory body can move fast enough to protect the world's physical archives before they are reduced to pulped training data remains an open question - a race against the shredder.