The KNAW
Corpus Creation and Automatic Alignment of Historical Dutch Dialect Speech
Pages
10
Time to read
31 mins
Publication
Language
English
Pages
10
Time to read
31 mins
Publication
Language
English
This technical report details the creation of a corpus that includes audio recordings and transcriptions of historical Dutch dialect speech. The corpus is based on recordings from the Dutch Dialect Database, which contains dialectal variations recorded across the Netherlands during the latter half of the twentieth century. The report outlines the methodology employed for aligning these audio recordings with their corresponding transcriptions and metadata. It discusses the challenges faced in automatic speech recognition due to the non-standard nature of historical transcriptions and the importance of digitization in making dialectal speech resources more accessible. The authors present their approach to aligning the recordings with transcriptions, which is designed to be generalizable to other language variants. The report also reviews the history of the Meertens Institute's dialect recordings and evaluates the effectiveness of the alignment method, concluding with recommendations for future research in dialect corpus building.