Science and Research Content

HathiTrust releases significantly expanded open dataset, the HathiTrust Research Center (HTRC) Extracted Features (EF) Dataset, Version 1.0 -

https://www.hathitrust.orgHathiTrust has announced the release of a significantly expanded open dataset, the HathiTrust Research Center (HTRC) Extracted Features (EF) Dataset, Version 1.0. This dataset provides researchers with open access to data extracted from the full text of the HathiTrust Digital Library (HTDL) at an unprecedented scale.

The Extracted Features Dataset opens the complete HathiTrust collection for investigations into historical and cultural trends, the rise and fall of topics within the corpus, and the evolution of words and writing structures in publications dating from the 16th to the late 20th century. The EF Dataset provides quantitative information about word and line counts, parts of speech, and other details within each page of every volume in the HTDL. In addition to these larger-scale investigations, the EF Dataset also allows researchers to closely analyze the contents of a given volume or subset of volumes.

The data is extracted from 13.7 million volumes found in the HTDL, representing over 5 billion pages consisting of over 2 trillion tokens (words). A preliminary release of the EF Dataset, drawn from a much smaller subset comprising only HathiTrust’s public domain collection, has already enabled novel research from scholars in economics, history, linguistics, literary studies and sociology, among other fields.

Founded in 2008 and hosted at the University of Michigan, HathiTrust preserves and provides access to millions of digitized books and journals from the collections of more than 120 institutional academic and research partners via its certified trusted digital repository. This searchable archive of published literature from around the world includes both in-copyright and public domain materials from mass digitization programs and partners’ local digitization initiatives.

The HathiTrust Research Center is an advanced research service of HathiTrust and a collaborative research center launched jointly by Indiana University and the University of Illinois. The Research Center team strives to meet the technical challenges that researchers face when dealing with massive amounts of digital text, by developing cutting-edge software tools and cyber infrastructure to enable advanced computational access to the growing digital record of human knowledge.

More information about the Extracted Features Dataset and access to it, go to https://analytics.hathitrust.org/datasets. The HTRC EF Dataset is released under a Creative Commons CC-BY license.

Brought to you by Scope e-Knowledge Center, a world-leading provider of metadata services, abstraction, indexing, entity extraction and knowledge organisation models (Taxonomies, Thesauri and Ontologies).

Click here to read the original press release.

STORY TOOLS

  • |
  • |

sponsor links

For banner ads click here