The majority of the fifth floor of the Van Pelt-Dietrich Library Center is currently closed for construction. Learn more about this renovation to the Zilberman Family Center for Global Collections, including how to request materials.

The RDDS Blog

When AI Meets Ancient Handwriting: Penn Libraries at SCOOP 2026

Penn Libraries staff from Cultural Heritage Computing and Research Data & Digital Scholarship traveled to the University of Vienna in Austria for the 2026 SCOOP Network conference, joining an international community of technologists, librarians, and scholars exploring how AI can help make handwritten texts and the cultures they represent more accessible.

Keynote speakers Anna Dolganov and David Smith demonstrate Apollo, a new AI model designed to transcribe and reconstruct damaged and fragmentary Ancient Greek texts.

Earlier this month, dozens of scholars working at the intersection of artificial intelligence, manuscript studies, and computational paleography gathered at the second annual conference for the Source Codes of the Past (SCOOP) Network, generously hosted by the University of Vienna in Austria. SCOOP Network’s stated mission is “making the handwritten heritage of the past machine-readable, across scripts, languages and centuries.” Penn Libraries staff on the Cultural Heritage Computing (CHC) and Research Data & Digital Scholarship (RDDS) teams were delighted to contribute to the conference from September 7–9, 2026.

Several scholarly debates emerged at SCOOP over how best to apply AI methods to solve the challenge of using machines to automate the transcription of ancient and medieval documents, which is enormously difficult in the field of computer vision due to the sheer number of languages and different handwritten scripts involved. The main debate concerned whether it is more effective to use an all-in-one, generalist AI model to transcribe documents in a single shot or else to employ multi-stage pipelines wherein preprocessing, AI transcription, and post-correction comprise separate steps. A second discussion concerned the sometimes-divergent aims of researchers and digital archivists, as researchers are often eager to study transcribed text even if it is imperfect, while archivists and librarians must ensure transcription accuracy and preserve the data for long-term use. These debates were conducted in a congenial scholarly spirit, with participants recognizing that closer alignment will accelerate progress toward their shared goals.

RDDS and the Penn Libraries hope to use insights gained from the SCOOP Network conference to improve their ongoing initiative to train new AI models for Handwritten Text Recognition, or HTR. As described by Penn’s Manuscript Collections As Data Research Group, “HTR uses AI to transcribe handwritten texts and can make corpora of manuscript materials searchable, help us discover unique texts, and turn handwritten texts into data for linguistic, historical, and quantitative analyses.” Especially useful, therefore, was the conference’s focus on what are sometimes called “digitally disadvantaged languages,” meaning languages and writing systems for which comparatively little digitized material or machine-readable training data is available. The conference highlighted how researchers are addressing these disparities by developing new datasets and adapting AI models to a growing range of scripts, including Ancient Greek, Icelandic, Arabic, Persian, Syriac, Armenian, Yiddish, Korean, Latin, Old Czech, Croatian Church Slavonic, Celtic languages, Glagolitic, Old Cyrillic, and Bhujimol, among others.

Penn Libraries staff Doug Emery and Jessie Dummer presented on ongoing efforts to bring HTR into library research and digital workflows. Emery, who serves as the Libraries’ Special Collections Digital Content Programmer, spoke on the topic of “Piloting HTR in the Library: Experiments and Groundwork,” describing work since 2023 to test HTR models on Penn’s manuscript collections and develop practical workflows for their use. Dummer, who serves as the Libraries’ Digitization Project Coordinator, presented on “Integrating HTR into Digital Library Workflows,” which focused on the challenges of moving from HTR experimentation to wider implementation, including discussion of a current project involving fifteenth-century Italian manuscripts led by Laura Morreale, who is the Schoenberg Institute for Manuscript Studies 2026–2027 Manuscript Collections As Data Fellow. Drawing on lessons from this project and others, Dummer explored how HTR can be incorporated into existing digital library workflows and how the resulting transcription data might be structured for preservation, reuse, and future research.

RDDS’s AI & Machine Learning Developer Christopher Demers Jimenez, Ph.D. also attended SCOOP to learn new approaches to transcribing Penn Libraries’ manuscript collections at scale. His participation will help RDDS build on the experimental work already underway at Penn and explore how emerging AI transcription methods can be applied responsibly and effectively across millions of pages of handwritten text in dozens of different languages. For instance, RDDS’s Data Science & Society Research Assistant Yifang Xia has previously worked on an HTR project that involved segmenting Ancient Chinese-Japanese texts, and RDDS’s Applied Data Science Librarian Jajwalya Karajgikar has similarly worked on HTR for South Asian scripts ranging from Sanskrit and Urdu to Tamil and Persian. Beyond work on manuscripts for well-resourced languages, RDDS thus aspires to expand its HTR work with potential new projects on Penn’s collection of Chinese rare book manuscripts, its large collection of Sanskrit documents, and its extensive Judaic Studies holdings in collaboration with the Judaica Digital Humanities initiative at the Kislak Center for Special Collections, Rare Books and Manuscripts.

Together, these efforts will inform the Libraries’ next steps in developing HTR infrastructure that can make Penn’s handwritten collections more accessible for research, discovery, and computational analysis.

Maps and More

Campus Libraries Map

Staff Information

Resources for Staff Committees