With a new agreement, the Penn Libraries expands access to HathiTrust text data
Our new agreement greatly expands the availability of HathiTrust's public domain and Creative Commons licensed text data for Penn researchers interested in computational analysis.
For nearly 20 years, HathiTrust has worked with libraries across the country to preserve and provide access to a wide variety of digitized books and other materials to researchers and other users. Penn Libraries has been a member library with HathiTrust for since 2010, and now a new agreement gives our researchers access to even more of what HathiTrust makes available.
Founded in 2008, HathiTrust is a collaborative digital repository with over 200 academic and research library members. Working with Google to digitize texts from various member libraries, HathiTrust preserves over 19 million books and other materials like magazines, newspapers, sheet music, journals, and government documents. It provides full-text access to as many of those works as it can under U.S. copyright law and offers special access to users with visual and print disabilities. The size of the collection is continually growing, as new institutions join and contribute to the collection.
What does our new deal give you?
Not only can users read materials digitized by HathiTrust, but they can also create and make use of datasets based on these materials. Our new agreement greatly expands the availability of public domain and Creative Commons licensed text data for researchers interested in computational analysis, such as text and data mining. Previously, Penn researchers had access to about 800,000 volumes available for this type of analysis. By signing a distribution agreement with Google, researchers now have access to 6.6 million items.
There are almost 1,000 shared collections in HathiTrust for researchers to use as starting points. Created by individuals or institutions, these collections may be large or small, specific or quite broad, and cover topics ranging from historic cookbooks, to glaciology, to the works of Charles Dickens, to post-viral syndromes. There are also collections that include all materials in HathiTrust that were published in a particular year, making it easy to find items that recently entered the public domain. More importantly, under our new agreement, researchers can create and extract data from their own custom collections, or they can extract all 6.6 million items.
How can you access these datasets?
While this material is now available for members of the Penn community, users need to identify the data that they want and file a request outlining their project with HathiTrust to gain access. When the request is approved, users will also be asked to sign a researcher agreement, after which they’ll receive more information about retrieving the data. More information about the entire process can be found on HathiTrust’s Requesting and Using Research Datasets page.
HathiTrust is a multifaceted resource and this new access to its collections as data offers exciting research opportunities for the Penn community. If you have any questions about HathiTrust or any other resources that the Libraries offer, reach out to a librarian for help.
Date
September 1, 2026