Readability Metrics for Machine Translation in Dutch: Google vs. Azure & IBM

Publication date

2023-04-01

Authors

Toledo, C. vanISNI 0000000527855495
Schraagen, M.P.ISNI 0000000419454950
Dijk, F. vanISNI 0000000527813009
Brinkhuis, Matthieu J. S.ORCID 0000-0003-1054-6683ISNI 0000000419480083
Spruit, MarcoISNI 0000000077172004

Editors

Advisors

Supervisors

Document Type

Article
Open Access logo

License

cc_by

Abstract

This paper introduces a novel method to predict when a Google translation is better than other machine translations (MT) in Dutch. Instead of considering fidelity, this approach considers fluency and readability indicators for when Google ranked best. This research explores an alternative approach in the field of quality estimation. The paper contributes by publishing a dataset with sentences from English to Dutch, with human-made classifications on a best-worst scale. Logistic regression shows a correlation between T-Scan output, such as readability measurements like lemma frequencies, and when Google translation was better than Azure and IBM. The last part of the results section shows the prediction possibilities. First by logistic regression and second by a generated automated machine learning model. Respectively, they have an accuracy of 0.59 and 0.61.

Keywords

English to Dutch quality estimation, Machine translation, Quality estimation, Squad 2.0

Citation

Toledo, C V, Schraagen, M, Dijk, F V, Brinkhuis, M & Spruit, M 2023, 'Readability Metrics for Machine Translation in Dutch: Google vs. Azure & IBM', Applied Sciences, vol. 13, no. 7, 4444, pp. 1-14. https://doi.org/10.3390/app13074444