Historic academic libraries have always been more than repositories of books. They preserve manuscripts, maps, theses, archival records, catalogues and other forms of intellectual heritage accumulated over centuries. In the age of artificial intelligence, however, these collections are entering a new phase. Digitisation, machine-readable metadata, optical character recognition, handwriting recognition and generative AI are changing how researchers discover and use historical materials.
The University of Oxford offers an important example of this transformation. Its Bodleian Libraries, founded in 1602, are combining long-established preservation responsibilities with experiments in large-scale digitisation and AI-assisted discovery. This does not simply mean putting old books online. It raises deeper questions about how historical collections can be converted into usable research data, how AI should assist librarians, and how universities should balance openness, intellectual property, accuracy and responsible technology use.
From Historical Collections to Digital Research Resources
Digitisation has been part of academic library work for decades. A manuscript that once required a researcher to travel physically to a reading room can now potentially be studied from another country.
The scale of the challenge, however, remains enormous.
Oxford's Bodleian Libraries hold approximately 14 million printed items. Digitising collections of this size requires far more than scanning individual pages. Institutions must determine which materials should be prioritised, what equipment should be used, how metadata should be created, how files should be preserved and how researchers will eventually discover the material.
This is where artificial intelligence is becoming increasingly relevant.
AI tools can potentially help extract text from documents, classify information, improve catalogue records and make collections searchable in ways that traditional catalogues cannot easily support.
Oxford's Bodleian Digitisation Research Project
In February 2025, the Bodleian Libraries began a pilot research project investigating how parts of their collections could be digitised at greater scale. The project is funded by OpenAI as part of a five-year collaboration with the University of Oxford and the broader NextGenAI higher-education initiative.
The project is examining several questions.
Can digitisation throughput be increased?
Can generative AI improve metadata and full-text records?
Which parts of enormous historical collections should be prioritised for digitisation?
Can digitisation take place efficiently outside conventional imaging studios?
And how might AI change the ways researchers search across library collections?
These questions illustrate an important shift. Libraries are no longer considering digitisation only as an image-capture activity. They are increasingly studying the entire chain from physical object to digital file, structured metadata, searchable text and research discovery.
Historical Dissertations as an AI Test Case
One of the collections being used in Oxford's experiments is the Bodleian's Global Dissertations collection.
The wider Oxford-OpenAI collaboration announced plans to digitise around 3,500 public-domain dissertations dating from 1498 to 1884. The intention is to make previously inaccessible material searchable and available more widely to students and researchers.
The Bodleian's digitisation research has also used its broader collection of nineteenth- and twentieth-century European and American dissertations as experimental material for metadata and transcription work.
Approximately 125,000 dissertation images were captured as part of this research, while another 429,000 files were created from scanned catalogue cards.
This demonstrates why historical libraries are increasingly important to AI-related research.
Their collections contain enormous quantities of structured and semi-structured knowledge that have never been available in a form suitable for large-scale computational analysis.
AI, OCR and Handwritten Records
A major challenge in historical digitisation is converting scanned pages into usable text.
Conventional Optical Character Recognition, or OCR, has long been used to convert printed images into machine-readable text. However, older documents present special problems. Historical fonts, damaged paper, unusual layouts and inconsistent printing can reduce OCR accuracy.
Handwritten catalogues are even more difficult.
Oxford researchers are therefore examining both OCR and handwriting text recognition technologies and comparing traditional methods with newer AI-enhanced approaches. The project is investigating whether AI can extract information from catalogue cards, classify elements of catalogue data and potentially assist in generating catalogue records directly from documents.
If these technologies become reliable enough, millions of historical pages could become considerably easier to search.
A researcher studying nineteenth-century medicine, for example, might eventually search across thousands of dissertations that previously existed only as physical documents or basic catalogue entries.
Librarians Remain Central
AI-assisted cataloguing does not necessarily mean removing librarians from the process.
Oxford's project explicitly examines how specialist librarians can remain humans in the loop in automated and semi-automated workflows.
This is essential because AI systems can make mistakes.
A model might misread a handwritten name, confuse publication dates or incorrectly classify a historical subject. Such errors can become particularly problematic when inaccurate metadata is reproduced across digital systems.
Professional librarians bring contextual knowledge, cataloguing standards and subject expertise that automated systems do not consistently possess.
The emerging model is therefore likely to involve AI increasing the speed of repetitive tasks while human specialists verify, interpret and govern the resulting information.
Collections Are Becoming Data
Another important development is the concept of collections as data.
Traditionally, researchers might retrieve individual books or manuscripts and read them one by one. Digitised collections can instead be made available in formats that permit computational analysis.
Oxford and the Staatsbibliothek zu Berlin have been collaborating on a project examining how cultural-heritage collections can be published as data and used by different research communities. The project notes that the rise of AI has increased demand for large-scale cultural-heritage datasets, while also raising challenges involving sensitive or contested historical material.
This could transform research in history, literature, linguistics and the humanities.
Researchers might use computational methods to identify linguistic changes across centuries, analyse networks of authors, examine historical geographic references or compare themes across thousands of texts.
AI therefore potentially changes not only access to collections but also the scale at which humanities research can be conducted.
AI Training Is Becoming Part of Library Education
Libraries are also becoming important centres for AI literacy.
Oxford states that its AI training programme has expanded so that staff and students can access ongoing support from units including its Digital Capabilities team, AI Competency Centre, Centre for Teaching and Learning and Bodleian Libraries.
From the 2025–26 academic year, Oxford also began providing staff and students with access to ChatGPT Edu alongside other AI tools, accompanied by university guidance and training around responsible use.
This represents a broader change in the academic role of libraries.
Historically, librarians trained students to search catalogues, use databases, evaluate sources and manage references. Those information-literacy functions remain essential, but they are now expanding to include AI literacy.
Researchers increasingly need to understand how AI-generated outputs are produced, when they require verification, how sensitive data should be handled and how AI tools interact with copyright and research ethics.
Training for the Wider Library Sector
Oxford's role also extends beyond its own students.
The Bodleian's Towards Digital Collections initiative is developing training resources for professionals working in galleries, libraries, archives and museums. The project, running from 2026 to 2028, is intended to strengthen digital-collections knowledge across the cultural-heritage sector.
The Digital Humanities at Oxford Summer School also includes training on AI in research libraries. Its learning objectives include understanding the strategic role of AI in cultural-heritage institutions, developing practical competence with AI tools and recognising ethical and responsible AI principles.
This suggests that AI literacy is becoming a professional requirement not only for computer scientists but also for archivists, librarians, historians and humanities researchers.
Copyright and Ethical Questions
The transformation of historical collections into AI-ready data also creates difficult legal and ethical questions.
Public-domain material is comparatively straightforward, but libraries also hold copyrighted works, personal correspondence, culturally sensitive material and collections involving contested histories.
The Bodleian itself has raised questions about how libraries should balance open-access goals with authors' intellectual-property rights and whether collection data should be supplied freely for commercial AI use.
These questions are likely to become increasingly important.
Digitising an item for scholarly access is not automatically the same as licensing it for machine-learning development.
Universities will therefore need clear policies concerning copyright, licensing, provenance, privacy and commercial reuse.
Risks of AI-Generated Metadata
Another concern is reliability.
AI-generated metadata may appear authoritative even when it is wrong.
A historical name may be incorrectly transcribed. An ambiguous date may be transformed into a false certainty. A language model may infer a subject category that was never present in the original source.
For this reason, historical AI projects require strong provenance.
Researchers should be able to distinguish between original catalogue information, machine-generated transcription and human-corrected metadata.
AI should make collections easier to explore without obscuring where information came from.
New Possibilities for Global Research
Despite these challenges, the potential benefits are substantial.
Large-scale digitisation can reduce geographical barriers to scholarship. Researchers who cannot travel to Oxford may gain access to sources previously available only in physical reading rooms.
Searchable historical texts can also support interdisciplinary research linking history with linguistics, computer science, geography, sociology and digital humanities.
Collections that have been overlooked because their catalogues were difficult to search may acquire new research visibility.
This is particularly important for international scholarship. Digitised academic collections can enable researchers across the world to engage with materials that historically required substantial travel funding and institutional access.
The Library of the AI Era
The emergence of AI does not make historical libraries less important.
It may make them more important.
AI systems require reliable knowledge sources, carefully maintained metadata and trustworthy provenance. Libraries have centuries of experience in precisely these activities.
The future academic library may therefore combine traditional preservation with digitisation laboratories, data services, AI-assisted discovery, computational research and digital-skills training.
Oxford's current work reflects this direction. The Bodleian Libraries' strategy already includes expanding scalable digitisation and developing expertise in AI and automated processing for the acquisition, cataloguing, preservation and use of special collections.
Conclusion
Oxford's experiments with historical collections illustrate a wider transformation taking place across academic libraries.
Books, dissertations, catalogue cards and archival materials are increasingly being converted into digital objects that can be searched, analysed and potentially used in AI-assisted research. At the same time, libraries are becoming places where students and researchers learn how to use artificial intelligence responsibly.
The transition is not simply technological. It raises fundamental questions about accuracy, copyright, intellectual property, cultural sensitivity, transparency and human oversight.
The most successful academic libraries of the AI era are therefore unlikely to abandon their traditional mission. Instead, they will extend it.
Their task will remain to preserve and provide access to knowledge—but increasingly that responsibility will include determining how historical knowledge is digitised, structured, interpreted and responsibly connected to artificial intelligence.