Detecting semantic overlap : A parallel monolingual treebank for Dutch
Files
Publication date
2008-11
Authors
Marsi, Erwin
Krahmer, Emiel
Editors
Advisors
Supervisors
DOI
Document Type
Part of book or chapter of book
Metadata
Show full item recordCollections
License
Abstract
This paper describes an ongoing effort to build a large-scale monolingual treebank of parallel/
comparable Dutch text, where nodes of syntax trees are aligned and labeled according to
a small set of semantic similarity relations. Such a corpus has many potential uses and applications
in e.g., multi-document summarization, question-answering and paraphrase extraction.
We describe the text material, preprocessing, annotation, and alignment of sentences
and syntax trees, both manual and automatic. Two new annotation tools are presented, as
well as results from pilot experiments on inter-annotator agreement. On the basis of this
resource, new automatic alignment software and NLP applications will be developed.