Can We Survive without Labelled Data in NLP? Transfer Learning for Open Information Extraction

Publication date

2020-09-01

Authors

Sarhan, IngyISNI 000000049306204X
Spruit, MarcoISNI 0000000077172004
Esposito, M.

Editors

Massala, G.
Minutolo, A.
Pota, M.

Advisors

Supervisors

Document Type

Article
Open Access logo

License

cc_by

Abstract

Various tasks in natural language processing (NLP) suffer from lack of labelled training data, which deep neural networks are hungry for. In this paper, we relied upon features learned to generate relation triples from the open information extraction (OIE) task. First, we studied how transferable these features are from one OIE domain to another, such as from a news domain to a bio-medical domain. Second, we analyzed their transferability to a semantically related NLP task, namely, relation extraction (RE). We thereby contribute to answering the question: can OIE help us achieve adequate NLP performance without labelled data? Our results showed comparable performance when using inductive transfer learning in both experiments by relying on a very small amount of the target data, wherein promising results were achieved. When transferring to the OIE bio-medical domain, we achieved an F-measure of 78.0%, only 1% lower when compared to traditional learning. Additionally, transferring to RE using an inductive approach scored an F-measure of 67.2%, which was 3.8% lower than training and testing on the same task. Hereby, our analysis shows that OIE can act as a reliable source task.

Keywords

Transfer learning, Open information extraction, Relation extraction, Recurrent neutral networks, Word embeddings

Citation

Sarhan, I, Spruit, M, Esposito, M, Massala, G (ed.), Minutolo, A (ed.) & Pota, M (ed.) 2020, 'Can We Survive without Labelled Data in NLP? Transfer Learning for Open Information Extraction', Applied Sciences, vol. 10, no. 17, 5758. https://doi.org/10.3390/app10175758