20 September, 2026
The first two posts in this series framed the “tiger eating its tail” dynamic of AI doom loops and mapped the “buy → despine → scan → discard” pipeline for rare and out‑of‑print books. This third post focuses on the legal and copyright fault lines: who owns these texts, on what basis AI companies use them, and where the law may yet draw a line.
| medievalist.net |
Out of print, but not out of copyright
A large share of the books targeted for AI training are:
- Out of print: no longer commercially available from publishers.
- Still under copyright: often published in the mid‑ to late‑20th century, well within current copyright terms.
That creates a tension: these works are hard to obtain legally at scale, yet highly valuable as clean, pre‑AI training data. AI companies and their vendors navigate this space through a mix of strategies:
- Relying on fair use or equivalent exceptions for text and data mining.
- Negotiating licenses or bulk agreements with rights holders or intermediaries.
- Operating in legal grey zones where enforcement is weak or slow.
Fair use and text‑and‑data‑mining exceptions
United States: fair use
In the U.S., AI companies often lean on fair use doctrine, arguing that:
- Training models is transformative: the system learns patterns, not reproducing works verbatim.
- Only small, non‑expressive fragments are used at any one time.
- The use does not directly substitute for the original market for the books.
Critics counter that:
- Mass ingestion of entire works, even if not directly output, still exploits the creative labor in those works.
- Commercial AI systems may compete with or undermine potential licensing markets for text and data mining.
- Courts have not yet fully settled how fair use applies to large‑scale AI training, especially for in‑copyright, out‑of‑print books.
Europe: text‑and‑data‑mining (TDM) exceptions
In the EU, the Copyright Directive provides specific text‑and‑data‑mining exceptions:
- A general TDM exception for research organizations and cultural heritage institutions.
- A broader TDM exception that can apply to others, but with more conditions and the possibility for rights holders to opt out.
How these exceptions interact with commercial AI training, and with books obtained via bulk purchase and destruction of originals, is still being tested in policy and, increasingly, in court.
Cultural‑heritage and export controls
Beyond copyright, some books are protected as cultural assets:
- National laws may restrict export of certain rare or historically significant volumes.
- Archives and special collections may be subject to specific rules about reproduction and destruction.
When rare books cross borders in bulk to reach scanning facilities, they can brush up against these rules, even if individual titles do not appear “significant” on paper.
Where the law may move next
Several developments could reshape this landscape:
-
Court rulings on AI training
Cases involving news archives, books, and image datasets will clarify how far fair use or TDM exceptions stretch. -
AI‑specific regulation
Emerging frameworks may require more transparency about training data sources, impact assessments, or respect for cultural and scholarly interests. -
Collective licensing models
Rights‑holder organizations may push for collective licensing schemes for AI training, similar to existing models for music or photocopying.
The legal question is not only “can they do this?” but “on what terms, and with what obligations?” The answer will help determine whether the tiger’s tail remains part of a shared cultural heritage, or becomes private fuel for a self‑consuming AI loop.
Author's note: a challenge to US "fair use" is shaping up with New York Times suing Microsoft & Open AI over the use of copyright work (New York Times - arrow back on browser or click Continue in tiny at bottom to read not subscribe).
No comments:
Post a Comment