19 September, 2026
![]() |
| Mouseion Alexandreias, "Shrine of the Muses of Alexandria" |
Reported cases and channels
While full transparency is rare, several recurring patterns have emerged from investigative reporting, librarian networks, and digital‑humanities discussions:
-
U.S.‑based AI labs and data vendors
Major American AI companies work through specialized data‑collection firms to acquire “long‑tail” text. These firms engage scanning contractors that explicitly handle high‑volume book workflows, including despine scanning. Corporate registry searches and procurement documents sometimes link these vendors back to known AI‑lab supply chains. -
Used‑book wholesalers shifting to bulk AI‑oriented sales
Dealers that once supplied libraries and secondhand shops now move entire collections in large lots, often with minimal public bibliographic detail. Some cite “client confidentiality” when asked about end users. -
Library deaccessioning pipelines
Academic and public libraries, under budget pressure, sell low‑circulation monographs in bulk. In some documented cases, institutions believed books were destined for “preservation digitization,” only to learn later that physical copies were destroyed after scanning and no public digital edition was created.
Exact company names are often obscured by layers of contracting, but the structure is consistent: AI labs → data/collection vendors → scanning operators → physical destruction of originals.
Regional differences and fault lines
United States
- Large, decentralized library system with many independent academic and public libraries.
- Strong market for used and remaindered books; flexible deaccessioning policies in many institutions.
- Relatively permissive environment for private data collection, with limited specific regulation of AI training data sources so far.
Europe
- More centralized national and research libraries, with stronger traditions of legal deposit and preservation mandates.
- Emerging AI regulation (e.g., EU AI Act) that increasingly touches on data governance, transparency, and fundamental rights, though specifics on training corpora are still evolving.
- Greater public scrutiny of cultural‑heritage privatization, with active scholar and librarian networks pushing back against opaque digitization deals.
Other regions
- In many countries, national libraries and archives are key gatekeepers of older print collections, making large‑scale private acquisition harder without state involvement.
- Local book markets and legal frameworks vary widely, affecting how easily bulk purchases and cross‑border scanning operations can be organized.
Libraries, laws, and policy angles
Library policies and ethics
Many libraries operate under missions that emphasize:
- Long‑term preservation of cultural heritage
- Equitable access for researchers and the public
- Transparency about deaccessioning and digitization partnerships
The “despine‑and‑discard for private AI training” model conflicts with all three. This has sparked internal debates in library communities about:
- Tightening deaccessioning guidelines to require public‑access digitization or retention of originals
- Demanding contractual guarantees that scanned content will be at least partially accessible to researchers
- Refusing sales to buyers whose end use is known to be purely proprietary and non‑transparent
Legal and regulatory levers
Several legal and policy tools could shape this space:
-
Copyright and licensing
Even when books are out of print, many remain under copyright. How AI companies license or justify use of these texts (fair use, specific agreements, etc.) will be a key battleground. -
Cultural‑heritage and export rules
Some jurisdictions treat certain books and manuscripts as protected cultural assets, restricting bulk export or destruction. -
AI‑specific regulation
Emerging frameworks may require more transparency about training data sources, impact assessments for high‑risk systems, or obligations to respect cultural and scholarly interests.
Possible responses and alternatives
Breaking the “tiger‑and‑child” loop does not mean stopping AI development; it means reshaping how training data is sourced and governed. Some possible directions:
-
Public‑interest digitization
Expand and fund library‑led or consortium‑led digitization where scans are retained, originals preserved, and outputs at least partly open to research. -
Transparent data trusts
Create governed pools of high‑quality text (including older books) with clear rules on access, attribution, and use, rather than fully private corpora. -
Library and scholar safeguards
Adopt deaccessioning policies that forbid destruction of unique or scarce items for purely private digitization, and require transparency about buyers and intended use. -
Regulatory pressure
Use emerging AI and cultural‑heritage regulation to require disclosure of major training data sources and to protect certain classes of materials from privatized, destructive scanning.
The map is still being drawn. But the choices made now—by AI companies, libraries, regulators, and scholars—will determine whether the tiger and the child keep circling the tree until they turn to butter, or whether they find a way to step off the path and change the story.

No comments:
Post a Comment