Monday, 21 September 2026

When the map loses its footnotes: AI, rare books, and historical‑GIS research

HTML: Perplexity AI
Source: what is doom loop
21 September, 2026


Update: please find at bottom a link to an in-depth curriculum from Spatial Reserves

Fourth in a series (last here) reviewing trials & tribulations in AI today. Fifth & last one follows here.

The first three posts in this series traced the AI doom loop, the rare‑book “buy → despine → scan → discard” pipeline, and the legal and copyright fault lines. This fourth post brings the story down to the level of historical and geospatial research: what happens to projects that rely on old maps, local histories, and specialized monographs when those sources are privatized, degraded, or made less accessible.

AI-generated by author (Medium)


Why historical‑GIS depends on “boring” books

Work in historical GIS, spatial history, and local heritage often leans on:

  • Old gazetteers and place‑name dictionaries that document vanished settlements, changed borders, and local variants.
  • Regional and local histories with detailed descriptions of land use, transport routes, and administrative boundaries.
  • Specialized monographs in archaeology, agrarian history, and historical geography that never made it into mainstream digital collections.

These are exactly the kinds of long‑tail, pre‑digital texts now being targeted for bulk acquisition and private AI training.


Risks for research workflows

Loss of physical copies and marginalia

When rare or unique volumes are despined and discarded:

  • Handwritten notes, ownership marks, and inserted materials disappear.
  • Binding structures and annotations that help date or contextualize a text are lost.
  • Future scholars cannot re‑examine the physical artifact with better imaging or new questions.

For historical‑GIS, where provenance and context often matter as much as the text itself, this is more than an abstract loss.

Privatization of key references

If the only high‑quality digital versions of certain books exist inside private AI training corpora:

  • Researchers cannot cite or verify passages directly from a stable, accessible edition.
  • AI tools may “know” content that humans cannot easily check or reference.
  • Reproducibility suffers: two researchers using different models may get different “facts” from the same underlying source.

Distortion through AI summarization

As historians and geographers increasingly use AI to:

  • Summarize old texts
  • Extract place names and dates
  • Generate candidate entries for gazetteers or attribute tables

there is a risk that errors or simplifications in model outputs become embedded in datasets and maps, then feed back into future models—a classic doom‑loop pattern, now in spatial form.


Opportunities, if governance is right

The same technologies can also help historical‑GIS work, if the underlying data is governed well:

  • Better OCR and layout analysis
    Modern models can dramatically improve transcription of old fonts, maps, and tables, if trained on diverse, high‑quality historical material that remains accessible.
  • Named‑entity and place‑linking
    AI can assist in linking historical place names to modern coordinates, but only if the source texts are citable and verifiable.
  • Cross‑collection discovery
    Well‑designed, transparent corpora can help researchers find relevant monographs, articles, and archival references across institutions and languages.

Practical stances for researchers and mappers

For those working in historical GIS and related fields, a few practical positions emerge:

  • Prefer open, citable digital editions
    When possible, base analyses on texts available through libraries, archives, or trusted repositories, not solely on AI‑generated summaries.
  • Document AI use explicitly
    In methods sections, note where AI was used for transcription, extraction, or interpretation, and keep raw outputs where feasible.
  • Support public‑interest digitization
    Advocate for library‑led or consortium‑led digitization that preserves originals and makes scans at least partially open to research.
  • Push for transparency
    Encourage vendors, publishers, and AI providers to disclose major training data sources, especially when they affect domain‑specific knowledge.


A map without reliable footnotes is dangerous; an AI‑assisted historical‑GIS workflow without transparent, accessible sources is worse. Keeping the tiger and the child from circling into butter means insisting that the texts behind our maps remain traceable, citable, and common.

Note: these practices also slot in with a last example on the main blog here, of AI complementing but not replacing GIS in a geo-history context.


Author's footnote

Spatial Reserves free online course: GIS and Society (web view)


“... [@josephkerski] rigorously tested, taught, refined, and evaluated this course in higher education full-semester settings, short courses, online and face-to-face, in a wide variety of settings for a wide diversity of GIS data users and analysts in multiple locations around the world.
This course invites learners to work with real data to solve problems and foster skills through a set of readings, hands-on activities, discussions, quizzes, and a final project.  I also invite instructors who seek to incorporate meaningful data and discussion about data in their courses to use all or part of this course in their own courses and programs.”

3 comments:

  1. Andrew - Many thanks for saying here what needs to be said, including "When possible, base analyses on texts available through libraries, archives, or trusted repositories, not solely on AI‑generated summaries"<< yes! We are not in the scholarly and/or GIS communities saying no to all AI generated summaries but if that is ALL we use, then... we are treading in dangerous waters. I think about historians as one of many examples who rely on finding notes in books they dig up in books in public libraries, and ticket stubs, old menus, and other treasures in physical form, that they use in their research and writing (such as David McCullogh's research and authoring). What can they use in the future if those bits are in summarized all-digital form, and even locked away in Facebook pages or some other zones where one might need an account to access? I also think about the classic ArcGIS Story Maps that went away in Feb 2026. I lost over 500 story maps in that transition; I had no staff to help me transition these documents, and many faculty and researchers were in the same situation and/or had no ability to run a Jupyter Notebook to convert these. It wasn't like MY story maps were anything STELLAR but I spent hundreds of hours compiling them; more importantly things like the Alan Lomax story maps of music in the Deep South from his interviews of musicians from 1979 - the blues, bluegrass, zydeco, and others - all vanished. I have been using GIS all my life and other tools as well and don't pine for paper-only clunky days, but we as a society need to think about when we transition the tools, what gets saved, and what vanishes? And what SHOULD be saved ? Thank you - Joseph Kerski

    ReplyDelete
  2. Thanks, on one hand McCollogh's "1776" that showed the Revolutionary War could've gone eitner way, on the other OpenAI's Navier-Stokes "proof" a perfect non-seqitur.

    And you know why I lost my story-maps? Gov.uk played fast&loose w data freedom when protecting private interests... so I had to repost significant chunks (only for East Anglia at that) of corrected data that put storage fees beyond my reach eventually!

    ReplyDelete
    Replies
    1. https://www.perplexity.ai/search/cd96f35f-424b-4c11-b014-cabb3a9137a6

      Delete