Friday, 18 September 2026

“The tiger eating its tail”, or how to break out of the “AI slop cycle”

HTML: Perplexity AI
Source: What is a doom loop
18 September, 2026

This follows on a history note here if we remember "whence AI?" First of a series continuing here.

Original tale, YouTube


There is an Indian folk image, often told as a kind of nursery rhyme, of a child and a tiger chasing each other around a tree, round and round, until both are worn down and “turn into butter.” It is a vivid picture of a futile, self‑consuming loop: motion without progress, effort that only exhausts and dissolves what it started with.

That image fits the emerging “AI slop cycle” uncannily well. As models train increasingly on text generated by earlier models, and as the public web fills with AI‑influenced content, we risk a kind of cultural churning: language and knowledge going in circles, losing texture and truth, until everything starts to feel smoothed into an indistinct paste.


Part I — Framing: AI doom loops and the rare‑book grab

A “doom loop” in AI is a self‑reinforcing cycle in which models increasingly train on their own outputs or on data shaped by earlier models, leading to loss of diversity, accuracy, and reliability. As the public web fills with AI‑generated text, the value of high‑quality, human‑authored, pre‑AI material rises sharply.

In that context, a quiet but significant trend has emerged: large AI companies, often through intermediaries, are buying up rare and out‑of‑print books, having them despined, scanned at high resolution, and then discarded. The physical volumes disappear; the text lives on inside private training corpora.

Why rare books matter for AI

  • They are pre‑digital and pre‑AI, offering clean, edited, human‑authored language.
  • They cover specialized domains (history, law, philosophy, science) that are underrepresented online.
  • They act as anchors against model collapse, helping to preserve linguistic and factual diversity.

From an AI‑safety and data‑strategy perspective, this is a deliberate countermeasure to the doom loop: secure high‑signal text before it is drowned out by synthetic content. But the practice raises deep questions about who controls cultural heritage, and what is lost when unique physical copies are turned into disposable inputs for proprietary models.


Part II — Inside the book‑buying pipeline for AI training

Investigations in 2024–2026 have uncovered a recurring workflow used by, or on behalf of, major AI companies:

  1. Bulk purchasing of older, often rare or out‑of‑print books through used‑book dealers, online marketplaces, and library deaccessioning sales.
  2. Intermediary scanning operations described as “digitization vendors” or “content preparation” firms, sometimes linked to AI labs’ broader vendor networks.
  3. Despining and high‑speed scanning: spines are cut so pages can be scanned flat; remaining paper blocks are then recycled or discarded.
  4. Private ingestion: text and images are added to proprietary training corpora, rarely released as public digital editions.

Who and how: documented actors and channels

Reporting points to several overlapping layers of actors:

  • Large AI labs and their data vendors
    Major U.S. AI companies contract specialized data‑collection firms to acquire “long‑tail” text underrepresented on the open web. Internal documents, vendor RFPs, and whistleblower accounts describe systematic acquisition of older monographs (pre‑2000, often pre‑1980), with emphasis on technical, scientific, legal, and humanities titles.
  • Used‑book wholesalers
    Dealers that normally supply libraries and secondhand shops have, in some cases, shifted to selling in very large lots (thousands of titles at a time) and stopped listing detailed bibliographic data publicly, citing “client confidentiality.”
  • Digitization and “content acquisition” firms
    Vendors advertising “full‑text extraction from print collections” sometimes mention “high‑volume book scanning for AI/ML clients” in marketing or recruitment materials. Some operate as shell LLCs or newly formed entities with minimal web presence, sharing addresses with known data‑labeling or AI‑infrastructure firms.
  • Library deaccessioning pipelines
    Academic and public libraries, under budget pressure, have deaccessioned low‑circulation monographs, including older scholarly titles, and sold them in bulk. In some documented cases, libraries believed books were going to “digitization for preservation,” only to discover later that physical copies had been destroyed after scanning, with no public digital version created.

What gets targeted

The books most in demand share a few traits:

  • Pre‑internet, pre‑AI publication dates (especially before 1990–2000).
  • Specialized and scholarly works in history, philosophy, law, economics, linguistics, and the sciences.
  • Out‑of‑print and rare editions with few surviving copies and limited presence in major public digitization projects.

These are exactly the texts that can help reduce reliance on noisy, AI‑contaminated web data and improve models’ performance on long‑tail, expert domains. In strategic terms, the pipeline functions as a clean‑data grab and a hedge against model collapse.

The costs

While this may help rein in aspects of the AI doom loop, the externalities are significant:

  • Cultural heritage loss: unique or scarce volumes are turned into disposable inputs for private models.
  • Scholarly access reduced: marginalia, bindings, and provenance marks disappear; future re‑scanning becomes impossible.
  • Privatization of knowledge: wealthy AI firms lock up chunks of the global knowledge commons in opaque, proprietary datasets.

The child and the tiger never escape the tree in the old tale; they simply wear themselves down. Our version of the story does not have to end that way. Breaking the “AI slop cycle” will require more than clever data acquisition: it will require new norms around cultural heritage, transparency about training corpora, and a shared commitment to keeping some part of the knowledge commons truly common.

No comments:

Post a Comment