Improving Successor Variety for Morphological Segmentation
Files
Publication date
2010-11
Authors
Çöltekin, Çağrı
Editors
Advisors
Supervisors
DOI
Document Type
Part of book or chapter of book
Metadata
Show full item recordCollections
License
Abstract
Successor variety is a commonly used measure for segmentation in language processing. It
is based on a simple idea that a large variety of letters (or phonemes) following an initial
word (or utterance) segment indicates a possible boundary. It dates back to Harris (1955),
and several methods based on successor variety have been used in the literature, particularly
for the purpose of segmenting words into morphemes. However, there have not been many
studies analyzing the measure itself. Even though the idea is simple and effective, the
current use in the literature does not utilize the measure to its full extent due to a number of
problems with the successor variety scores. This paper intends to address these problems by
introducing a normalization method, and demonstrates—using segmentation experiments
on two typologically different languages— the effectiveness of this improvement on the
morphological segmentation task.