-
Notifications
You must be signed in to change notification settings - Fork 1
Expand file tree
/
Copy pathabout page
More file actions
8 lines (6 loc) · 1.48 KB
/
Copy pathabout page
File metadata and controls
8 lines (6 loc) · 1.48 KB
1
2
3
4
5
6
7
8
Welcome to the Extracted Features Café (EFC)!
The EFC is a widget designed to make the HathiTrust Extracted Features API (https://htrc.stoplight.io/docs/ef-api/8xpvh96ani2e0-ef-api) accessible to nonprogrammers and to expand accessibility by allowing users without to search for the Extracted Features using common unique identifiers such as OCLC, ISSN, ISBN, LCCN, HTID, and HathiTrust record number. By providing user-friendly tools and data visualizations, the EFC aims to bridge the gap between humanities students and computational resources. This will empower them to leverage digital libraries without needing extensive programming knowledge.
What are Extracted Features?
The Extracted Features (https://htrc.github.io/torchlite-handbook/ef.html) are data elements derived from a chosen text in the HathiTrust Digital Library. The EFC presents these page- and volume-level features in a word frequency chart, word cloud, and frequency evolution chart, along with four word clusters based on topic modeling within the text. Word frequency charts exclude “stop words,” or common words such as articles, being verbs, prepositions, and pronouns from their results.
The EFC was created by Danielle Nasenbeny, Dolsy Smith, Eryclis Silva, and Lívia Clarete, during the Tools for Open Research and Computation with HathiTrust: Leveraging Intelligent Text Extraction (TORCHLITE) Hackathon hosted by the HathiTrust Research Center at University of Illinois at Urbana-Champaign.
Last Updated May 23, 2024.