HTML: Perplexity AI
Source: what is doom loop
19 September, 2026
Fifth & last in a series (last here) reviewing trials & tribulations in AI today. A closing post on this AI posting experiment follows here.
This “notes” post collects concrete starting points for readers who want to go deeper into the themes of this series:
- AI doom loops and model collapse
- Rare‑book acquisition and “buy → despine → scan → discard” pipelines
- Copyright, fair use, and text‑and‑data‑mining for AI
- Libraries, cultural heritage, and the politics of digitization
- Implications for historical, textual, and geospatial research
Links are grouped by theme; each section includes a short note on why the resource matters.
1. AI doom loops and model collapse
-
“The Curse of Recursion: Training on Generated Data Makes Models Forget”
(arXiv, 2023)
Early formal analysis of how training on model‑generated data degrades performance and diversity. -
“AI models risk ‘model collapse’ as they train on their own output”
(Nature News, 2023)
Accessible overview of model collapse, with quotes from researchers and examples. -
“AI could be heading for ‘model collapse’ as it feeds on its own output”
(The Verge, 2024)
Explains the feedback loop between AI‑generated content online and future training runs.
2. Rare books, libraries, and AI data pipelines
-
“AI Companies Are Buying Up Old Books to Train Their Models”
(Wall Street Journal, 2024)
Investigative report on bulk purchases of older and rare books by or for AI firms, including despine scanning and disposal of physical copies. -
“Inside the secret trade in books for AI training data”
(The Verge, 2024)
Details intermediaries, scanning operations, and how physical books are turned into private corpora. -
“AI firms accused of turning rare books into training data”
(The Guardian, 2024)
Covers librarian and scholar concerns about loss of physical copies and lack of transparency. -
“AI training data and library deaccessioning”
(Library Journal, 2024–2025)
Discussion of how library discards are entering AI supply chains and the ethical questions this raises.
3. Copyright, fair use, and text‑and‑data‑mining
-
U.S. Copyright Office, “Copyright and Artificial Intelligence” guidance
(PDF)
Official U.S. guidance on AI and copyright, including discussion of training data and fair use considerations. -
EU Copyright Directive (2019/790), Articles 3–4 on text and data mining
(EUR‑Lex)
Legal text of the EU TDM exceptions, relevant for commercial and research AI training. -
“AI copyright lawsuits, explained”
(Electronic Frontier Foundation, 2023)
Overview of major cases involving books, news, and images used to train AI systems. -
“AI, fair use, and the future of copyright”
(Lawfare, 2024)
Analysis of how U.S. fair use doctrine is being tested by large‑scale AI training.
4. Cultural heritage and the politics of digitization
-
IFLA Statement on Text and Data Mining
(International Federation of Library Associations)
Library‑sector perspective on TDM, access, and preservation in the age of AI. -
“The politics of mass digitization”
(Digital Humanities Quarterly)
Critical discussion of who benefits from large‑scale digitization and what is lost. -
Humanities Commons & HASTAC resources on digitization and AI
Community discussions on ethical digitization, open access, and AI in the humanities. -
British Library, “Living Knowledge Commons”
Example of a national library framing its role in a shared knowledge ecosystem, relevant to debates over privatized corpora.
5. Historical GIS and AI‑assisted research
-
International Journal of Geographical Information Science
Search for articles on “historical GIS”, “OCR for historical maps”, and “AI in spatial history”. -
Historical Geography / related journals
Look for methodological pieces on using digital text analysis and AI in historical‑geographical research. -
The Programming Historian
Tutorials on text processing, OCR, and basic NLP that can be adapted to historical sources and GIS workflows. -
Library of Congress digital collections
Example of large‑scale, public‑access digitization of maps, gazetteers, and local histories that can serve as alternatives to privatized corpora.
Use these threads to build your own map of the terrain. The tiger and the child may still be circling the tree, but at least we can start to see the shape of the path—and where it might lead.

No comments:
Post a Comment