Text Mining Methods for Automated Data Extraction from Health Technology Assessment Reports of Medicines Using Classical Natural Language Processing and Generative Artificial Intelligence

Publication date

2026-04-27

Authors

Versteeg, Jan-WillemISNI 0000000527858311
De Bruin, MariekeORCID 0000-0001-9197-7068ISNI 0000000397182332
Schermer, MaartenORCID 0000-0001-6770-3155
Najafabadi, Shiva NadiORCID 0009-0008-7308-1952
Mitra, ModhuritaORCID 0009-0008-0843-7537
Leopold, ChristineORCID 0000-0002-2046-8490ISNI 0000000512552117
Mantel - Teeuwisse, AukjeISNI 0000000390595150
Goettsch, W.G.ORCID 0000-0002-8022-7496ISNI 0000000395155859
Bloem, Lourens T.ORCID 0000-0002-0014-8625ISNI 000000049260699X

Editors

Advisors

Supervisors

Document Type

Article
Open Access logo

License

cc_by

Abstract

Objective This proof of concept for utilizing automatic data extraction methods to extract health technology assessment (HTA) attributes from HTA reports of medicines aimed to explore which attributes could be extracted and how accurately, using different data extraction methods. This enables easy access to insights into HTA recommendations for policymaking and policy-related research. Materials and Methods In total, 14 relevant attributes (eg, assessment outcome or date) were identified for extraction using two classical natural language processing (NLP) methods (rule-based and classification models) and a generative AI method (large language model (LLM)-based, i.e., Claude 3 Opus). The performance of these techniques was compared using 50 HTA reports published by the National Institute for Health and Care Excellence (NICE, United Kingdom). Results All three methods were able to extract certain attributes with high accuracy, with differences between the extraction methods and the type of attribute. The LLM-based extraction was the only method able to extract attributes on a medicine-indication combination level. The LLM-based extraction performed best (88–98% semantical accuracy for 12/14 attributes). Extraction of Outcome relative effectiveness analyses (REA) and Comparator was the most challenging and had the lowest accuracy (∼70% for the LLM-based extraction). Discussion & Conclusion Automatic data extraction for relevant attributes from HTA reports is possible, but there is still room for improvement. LLM-based extraction outperformed the two NLP methods, but challenges regarding the use of commercial software and reproducibility remain. Future research should focus on expanding the system to other HTA organizations and further refining the LLM-based extraction.

Keywords

Automated data extraction, generative AI, health technology assessment, large language models, natural language processing, Health Informatics

Citation

Versteeg, J-W, Bruin, M D, Schermer, M, Najafabadi, S N, Mitra, M, Leopold, C, Mantel-Teeuwisse, A, Goettsch, W & Bloem, L T 2026, 'Text Mining Methods for Automated Data Extraction from Health Technology Assessment Reports of Medicines Using Classical Natural Language Processing and Generative Artificial Intelligence', JAMIA Open, vol. 9, no. 2, ooag051. https://doi.org/10.1093/jamiaopen/ooag051