FIZ Karlsruhe
NLP Approaches for Enhancing Metadata Quality
Pages
13
Time to read
41 mins
Publication
Language
English
Pages
13
Time to read
41 mins
Publication
Language
English
This paper is a research article that discusses the challenges faced by cultural heritage institutions regarding the quality of bibliographical metadata collections, particularly those describing pre-modern objects. It outlines how incomplete and inaccurate metadata hampers the identification of literary works and complicates search and retrieval processes. The authors explore various natural language processing (NLP) approaches that leverage the length and descriptive nature of titles to improve metadata quality. The paper details the metadata collection of the Deutsche Digitale Bibliothek (DDB), which is encoded as linked open data and stored in a knowledge graph. It also reviews related literature on knowledge graph construction and information extraction, emphasizing the importance of accurately identifying bibliographic entities. The findings aim to assist librarians in enhancing their metadata practices, ultimately facilitating better content exploration and recommendations in digital libraries. The article concludes with a discussion on future work in this area.