A filter for syntactically incomparable parallel sentences

Publication date

2019-12

Authors

Kroon, Martin
Barbiers, L.C.J.ISNI 0000000106536912
Odijk, JanISNI 0000000397024625
Pas, Stéfanie van der

Editors

Berns, Janine
Tribushinina, Elena

Advisors

Supervisors

DOI

Document Type

Part of book
Open Access logo

License

taverne

Abstract

Massive automatic comparison of languages in parallel corpora will greatly speed up and enhance comparative syntactic research. Automatically extracting and mining syntactic differences from parallel corpora requires a pre-processing step that filters out sentence pairs that cannot be compared syntactically, for example because they involve “free” translations. In this paper we explore four possible filters: the Damerau-Levenshtein distance between POS-tags, the sentence-length ratio, the graph-edit distance between dependency parses, and a combination of the three in a logistic regression model. Results suggest that the dependency-parse filter is the most stable throughout language pairs, while the combination filter achieves the best results

Keywords

Taverne, Language and Linguistics, Artificial Intelligence

Citation

Kroon, M, Barbiers, S, Odijk, J & Pas, S V D 2019, A filter for syntactically incomparable parallel sentences. in J Berns & E Tribushinina (eds), Linguistics in the Netherlands 2019. AVT Publications, John Benjamins, Amsterdam, pp. 147-161.