Harnessing Textbooks for High-Quality Labeled Data: An Approach to Automatic Keyword Extraction

Publication date

2023-07

Authors

Pozzi, Lorenzo
Alpizar-Chacon, IsaacORCID 0000-0002-6931-9787ISNI 0000000506317436
Sosnovsky, SergeyISNI 0000000352729779

Editors

Advisors

Supervisors

DOI

Document Type

/dk/atira/pure/researchoutput/researchoutputtypes/contributiontojournal/conferencearticle
Open Access logo

License

cc_by

Abstract

As textbooks evolve into digital platforms, they open a world of opportunities for Artificial Intelligence in Education (AIED) research. This paper delves into the novel use of textbooks as a source of high-quality labeled data for automatic keyword extraction, demonstrating an affordable and efficient alternative to traditional methods. By utilizing the wealth of structured information provided in textbooks, we propose a methodology for annotating corpora across diverse domains, circumventing the costly and time-consuming process of manual data annotation. Our research presents a deep learning model based on Bidirectional Encoder Representations from Transformers (BERT) fine-tuned on this newly labeled dataset. This model is applied to keyword extraction tasks, with the model’s performance surpassing established baselines. We further analyze the transformation of BERT’s embedding space before and after the fine-tuning phase, illuminating how the model adapts to specific domain goals. Our findings substantiate textbooks as a resource-rich, untapped well of high-quality labeled data, underpinning their significant role in the AIED research landscape.

Keywords

automatic keyword extraction, BERT fine-tuning, labeled data, textbooks, General Computer Science

Citation

Pozzi, L, Alpizar-Chacon, I & Sosnovsky, S 2023, 'Harnessing Textbooks for High-Quality Labeled Data : An Approach to Automatic Keyword Extraction', CEUR Workshop Proceedings, vol. 3444, pp. 66-77.