Last week I attended the Conference of European National Librarians in Luxembourg. I came back with the feeling I had a lot of pieces of a puzzle, but without the image on the box. We talked about Europe, culture, AI, security, resilience, data, multilingualism and authorship. Last year I was inspired by CENL to start writing a series of essays on AI. To keep myself intellectually honest (it is very hard to unlearn something once you have learned it, so you mostly don’t even remember how it was not to know it) and to help me work through some of the fundamental questions that arise from it.
I have long used the concept of the creative communication cycle to create a context to understand human communication across time and space through text. In it readers’ and writers’ minds connect through text. So someone wants to know or read something and finds a text written by someone else. Writers create text for someone else to interact with. This used to be a fully human interaction, aided by processes of production, holding and hosting of the works, and access. In last year’s post, I used it to understand where AI was changing the landscape as I knew it.
Now I understand that there is a nuance missing in the model. The model implies people end up at sources, whether digital or physical. I already knew there was a difference between ‘current’ and ‘historical’ sources. These I learnt in my history study to call ‘primary’ and ‘secondary’ sources. Primary sources are original sources: government documents, diaries, photographs, contracts, eyewitness reports. Secondary sources interpret those records. These two types of sources have different values: the original sources can be re-interpreted according to current views; the secondary sources of one age become the primary sources of another.
I now realise that this model is no longer adequate, as AI is adding a new layer to this process. There is a knowledge infrastructure inherent in this way of humans encountering texts. Traditionally, tertiary sources played a role in this. A tertiary source is a source that organizes, summarizes, indexes, or compiles information from primary and secondary sources to help people find or understand existing knowledge. Think encyclopedias (including wikipedia), dictionaries, library catalogues. They are part of the infrastructure of knowledge.
Historically, every expansion of information came with new ways to order and use that information. Once the volume of knowledge exceeds what individuals can navigate directly, new information systems evolve to reduce cognitive complexity. Manuscripts came in relatively low number, and required a human scribe for every copy. Lists sufficed to keep track. The printing press made copies much cheaper and more abundant – a long period of distribution of paper copies followed and catalogues were created for access. In the seventies these catalogues (metadata) became digital, so these distributed collections became visible in networks. Then the information itself became digital, especially in research. Libraries started digitizing to make primary sources available online, thus keeping them part of the current conversation. Digital preservation of today’s primary sources started. Visualisations, ngrams and indexing for search engines became new tools.
The digital turn in society created massive amounts of data. This, combined with AI technologies that have been developed for decades led to foundation models. These are AI models trained on vast amounts of data to learn general patterns in language, images, audio, or other forms of information. Rather than being built for a single task, they provide a common foundation that can be adapted to many different uses, from chatbots and translation to image generation and scientific research. Large language models such as ChatGPT made this shift visible to the public in 2022 by making foundation models accessible through natural language conversation.
Rather than replacing tertiary sources, foundation models introduce a new kind of tertiary infrastructure. It is doing far more than just summarizing – it is generating text in response to questions, text that is probabilistic and personalised and varies by prompt, context (paid or not), and model version. It is no longer fixed and referential, meaning there is no stable relationship to any underlying document. It is as I argued elsewhere a synthetic reasoner. This means that how knowledge is formed (epistemic authority) is no longer visible. Ian Milligan recently called this a diminishing of the ‘Legibility of Absence’. When you visited one library, you were aware this is not all the knowledge in the world, when you use a search engine, this is harder, with AI you no longer see what is not there. So we need to reimagine the ecosystem of knowledge in the face of AI. This is a more fundamental shift than I had realized: it is the picture on the box.
Every generation has one or two technologies that fundamentally change how people interact with the world, not just what tools they use. From muscle to engine power, from physical messages to wire. These technologies change the default and require society to reorganize its infrastructure to accommodate it. In this generation – we have two major transformational technologies: digital and AI. These developments mean we live in a time of epistemic turbulence; rapid changes in how information is produced, distributed, and validated make it increasingly difficult for individuals and institutions to establish what is trustworthy, relevant, and true.
The word epistemic can feel exclusive: but it is helpful as a concept to define how we know what we know. Epistemic turbulence implies there is an ecosystem for information and knowledge and that the foundations of this ecosystem are challenged by new technology. Misinformation, AI slop, platform dependency, recalibration of narrative authority are more traditionally in the information domain. But now cyber-, climate and geopolitical risks are all factoring into this epistemic turbulence. The information domain has become much more political: Russia for instance describes this domain as information confrontation: a continuous struggle over information, perception, and decision-making. Cyber crime has already destroyed the digital ecosystem of the British Library. And climate change threatens physical and digital collections alike.
Mass literacy is a relatively recent phenomenon, people read for entertainment and when other forms of entertainment arrived, time spent reading for pleasure declined. In the same way, many people learnt how to use primary and secondary sources because they had to, not because it filled a need. Now that is no longer necessary, we need to think differently about how we as a society want our ecosystem of knowledge and information to work and incorporate the value of this work in new technologies, as well as preserving some of the older ones. For most people this will not be visible – they will accept the new synthetic reasoner, because being right most of the time is good enough. This feeling is not a failure to be corrected but a fact to work with.

Libraries are epistemic institutions – institutions that help create shared reference points for trust and meaning. In times of epistemic turbulence, they need to recalibrate. And I’d argue they have to recalibrate around people. Enhancing the ecosystem of knowledge, without aiming to remain the access point necessarily. This work will take many forms and libraries of course only have a contributing role to play in this great societal challenge. Education, research, culture to name but a few will also rethink their roles, as befits such a general purpose technology.
But for the knowledge ecosystem, I’d like to highlight a few examples I have encountered this year as ways for libraries to respond to the epistemic demands of this time. These are not answers by any means, but maybe beginnings of where to look going forward. Libraries should:
- Make knowledge structures legible again
Historian Jo Guldi has recently been working on computational history using multiple language editions of Wikipedia to understand differences in historical narratives. In this way showing where different language communities converge and diverge (link). This is in line with the work Marieke van Erp and Victor de Boer have been doing on bias and polyvocal knowledge graphs. At its core, this means making the knowledge embedded in texts explicit rather than implicit. By visualising how ideas, concepts and relationships are represented, underlying assumptions become visible and open to questioning, instead of being implicit and taken for granted.
2. Reduce structural bias in the infrastructure and build public alternatives.
AI, as I have argued in another essay, is making the world more unequal (at least in the short term). This is true socially as well as culturally. Martin Öövel, the national librarian of Estonia, for example, points to a Gini coefficient of 0.92 for the distribution of AI training data across languages, revealing how a small number of dominant languages disproportionately shape the knowledge embedded in foundation models. An initiative called the European Book Data Commons tries actively to bring together a large multilingual book data set for use in public AI. This will not solve the imbalance, but give public alternatives a chance.
3. Experiment with specialised services based on trustworthy sources.
This type of AI is called Retrieval-Augmented Generation (RAG), an AI approach that combines information retrieval with text generation. Before producing an answer, the model searches a specified collection of documents or databases and uses the retrieved information as evidence for its response. The Curator-Bot at the KB national library is an example, combining all knowledge on one seventeenth century work.
4. Ask their politicians to develop policy frameworks for the public access to information.
This means legislating for the public values included in library and media laws (trustworthiness, accessibility, independence, authenticity and pluralism) in the AI ecosystem. That requires people in all ministries, not just Home Office and Economic Affairs to be involved in Digital Affairs.
5. And last but not least: preserve access to the primary sources.
Primary sources remain the anchor of the knowledge ecosystem. Every generation reinterprets them differently, but they cannot be regenerated after they are lost. In the digital domain they are vulnerable as never before. The Global resilient information network, which Wilma van Wezenbeek presented at CENL is an emerging initiative around shared stewardship of digital assets of knowledge and culture.
And so a year on, I return to Richard Ovenden quoting Jacques Derrida: ‘There is no political power, without power over the archive’. Preserving the means to make sense of the world as we know it. Preserving the conditions under which future societies can still argue, remember, and imagine. Because people matter to people, people don’t matter to machines. Humans are still the level at which meaning is made.
June 2026
Other Essays: https://elsbethkwant.nl/essays-in-ai/