Saturday, 19 September 2026

Mapping the AI rare‑book pipeline: cases, regions, and policy fault lines

HTML: Perplexity AI
Source: what is doom loop
19 September, 2026

Second in a series (last herereviewing trials & tribulations in AI today. Third one follows here.

The previous post framed the “tiger eating its tail” dynamic: AI doom loops pushing companies to acquire pre‑digital, human‑authored text, and the resulting “buy → despine → scan → discard” pipeline for rare and out‑of‑print books. This follow‑up maps how that pattern shows up in different places, what roles libraries and laws play, and where pressure points for change might lie.

Mouseion Alexandreias, "Shrine of the Muses of Alexandria"


Reported cases and channels

While full transparency is rare, several recurring patterns have emerged from investigative reporting, librarian networks, and digital‑humanities discussions:

  • U.S.‑based AI labs and data vendors
    Major American AI companies work through specialized data‑collection firms to acquire “long‑tail” text. These firms engage scanning contractors that explicitly handle high‑volume book workflows, including despine scanning. Corporate registry searches and procurement documents sometimes link these vendors back to known AI‑lab supply chains.
  • Used‑book wholesalers shifting to bulk AI‑oriented sales
    Dealers that once supplied libraries and secondhand shops now move entire collections in large lots, often with minimal public bibliographic detail. Some cite “client confidentiality” when asked about end users.
  • Library deaccessioning pipelines
    Academic and public libraries, under budget pressure, sell low‑circulation monographs in bulk. In some documented cases, institutions believed books were destined for “preservation digitization,” only to learn later that physical copies were destroyed after scanning and no public digital edition was created.

Exact company names are often obscured by layers of contracting, but the structure is consistent: AI labs → data/collection vendors → scanning operators → physical destruction of originals.


Regional differences and fault lines

United States

  • Large, decentralized library system with many independent academic and public libraries.
  • Strong market for used and remaindered books; flexible deaccessioning policies in many institutions.
  • Relatively permissive environment for private data collection, with limited specific regulation of AI training data sources so far.

Europe

  • More centralized national and research libraries, with stronger traditions of legal deposit and preservation mandates.
  • Emerging AI regulation (e.g., EU AI Act) that increasingly touches on data governance, transparency, and fundamental rights, though specifics on training corpora are still evolving.
  • Greater public scrutiny of cultural‑heritage privatization, with active scholar and librarian networks pushing back against opaque digitization deals.

Other regions

  • In many countries, national libraries and archives are key gatekeepers of older print collections, making large‑scale private acquisition harder without state involvement.
  • Local book markets and legal frameworks vary widely, affecting how easily bulk purchases and cross‑border scanning operations can be organized.

Libraries, laws, and policy angles

Library policies and ethics

Many libraries operate under missions that emphasize:

  • Long‑term preservation of cultural heritage
  • Equitable access for researchers and the public
  • Transparency about deaccessioning and digitization partnerships

The “despine‑and‑discard for private AI training” model conflicts with all three. This has sparked internal debates in library communities about:

  • Tightening deaccessioning guidelines to require public‑access digitization or retention of originals
  • Demanding contractual guarantees that scanned content will be at least partially accessible to researchers
  • Refusing sales to buyers whose end use is known to be purely proprietary and non‑transparent

Legal and regulatory levers

Several legal and policy tools could shape this space:

  • Copyright and licensing
    Even when books are out of print, many remain under copyright. How AI companies license or justify use of these texts (fair use, specific agreements, etc.) will be a key battleground.
  • Cultural‑heritage and export rules
    Some jurisdictions treat certain books and manuscripts as protected cultural assets, restricting bulk export or destruction.
  • AI‑specific regulation
    Emerging frameworks may require more transparency about training data sources, impact assessments for high‑risk systems, or obligations to respect cultural and scholarly interests.

Possible responses and alternatives

Breaking the “tiger‑and‑child” loop does not mean stopping AI development; it means reshaping how training data is sourced and governed. Some possible directions:

  • Public‑interest digitization
    Expand and fund library‑led or consortium‑led digitization where scans are retained, originals preserved, and outputs at least partly open to research.
  • Transparent data trusts
    Create governed pools of high‑quality text (including older books) with clear rules on access, attribution, and use, rather than fully private corpora.
  • Library and scholar safeguards
    Adopt deaccessioning policies that forbid destruction of unique or scarce items for purely private digitization, and require transparency about buyers and intended use.
  • Regulatory pressure
    Use emerging AI and cultural‑heritage regulation to require disclosure of major training data sources and to protect certain classes of materials from privatized, destructive scanning.

The map is still being drawn. But the choices made now—by AI companies, libraries, regulators, and scholars—will determine whether the tiger and the child keep circling the tree until they turn to butter, or whether they find a way to step off the path and change the story.

 


No comments:

Post a Comment