Lessons from a User Experience Evaluation of NLP Interfaces

Publication date

2025

Authors

Calò, EduardoISNI 0000000512510558
Penkert, Lydia
Mahamood, Saad

Editors

Chiruzzo, Luis
Ritter, Alan
Wang, Lu

Advisors

Supervisors

Document Type

Part of book
Open Access logo

License

cc_by

Abstract

Human evaluations lay at the heart of evaluations within the field of Natural Language Processing (NLP). Seen as the “golden standard” of evaluations, questions are being asked on whether these evaluations are both reproducible and repeatable. One overlooked aspect is the design choices made by researchers when designing user interfaces (UIs). In this paper, four UIs used in past NLP human evaluations are assessed by UX experts, based on standardized human-centered interaction principles. Building on these insights, we derive several recommendations that the NLP community should apply when designing UIs, to enable more consistent human evaluation responses.

Keywords

Computer Networks and Communications, Hardware and Architecture, Information Systems, Software

Citation

Calò, E, Penkert, L & Mahamood, S 2025, Lessons from a User Experience Evaluation of NLP Interfaces. in L Chiruzzo, A Ritter & L Wang (eds), 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics : Proceedings of the Conference Findings, NAACL 2025. 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Proceedings of the Conference Findings, NAACL 2025, Association for Computational Linguistics (ACL), pp. 2915-2929, 2025 Annual Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics, NAACL 2025, Albuquerque, United States, 29/04/25. https://doi.org/10.18653/v1/2025.findings-naacl.159, conference