Corpus Creation and Automatic Alignment of Historical Dutch Dialect Speech
Publication date
2024-05
Editors
Calzolari, Nicoletta
Kan, Min-Yen
Hoste, Veronique
Lenci, Alessandro
Sakti, Sakriani
Xue, Nianwen
Advisors
Supervisors
DOI
Document Type
Part of book
Metadata
Show full item recordCollections
License
cc_by_nc
Abstract
The Dutch Dialect Database (also known as the 'Nederlandse Dialectenbank') contains dialectal variations of Dutch that were recorded all over the Netherlands in the second half of the twentieth century. A subset of these recordings of about 300 hours were enriched with manual orthographic transcriptions, using non-standard approximations of dialectal speech. In this paper we describe the creation of a corpus containing both the audio recordings and their corresponding transcriptions and focus on our method for aligning the recordings with the transcriptions and the metadata.
Keywords
corpus creation, dialectal speech, Dutch language variants, speech transcriptions, Theoretical Computer Science, Computational Theory and Mathematics, Computer Science Applications
Citation
Bentum, M, Sanders, E, van den Bosch, A, Zeldenrust, D & van den Heuvel, H 2024, Corpus Creation and Automatic Alignment of Historical Dutch Dialect Speech. in N Calzolari, M-Y Kan, V Hoste, A Lenci, S Sakti & N Xue (eds), 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings. 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC-COLING 2024 - Main Conference Proceedings, European Language Resources Association (ELRA), pp. 4021-4029, Joint 30th International Conference on Computational Linguistics and 14th International Conference on Language Resources and Evaluation, LREC-COLING 2024, Hybrid, Torino, Italy, 20/05/24. < https://aclanthology.org/2024.lrec-main.357 >, conference