Evaluating the construct validity of text embeddings with application to survey questions

Publication date

2022-12

Authors

Fang, QixiangORCID 0000-0003-2689-6653ISNI 0000000493063739
Nguyen, DongISNI 0000000419527451
Oberski, Daniel LeonardORCID 0000-0001-7467-2297ISNI 0000000396652603

Editors

Advisors

Supervisors

Document Type

Article
Open Access logo

License

cc_by

Abstract

Text embedding models from Natural Language Processing can map text data (e.g. words, sentences, documents) to meaningful numerical representations (a.k.a. text embeddings). While such models are increasingly applied in social science research, one important issue is often not addressed: the extent to which these embeddings are high-quality representations of the information needed to be encoded. We view this quality evaluation problem from a measurement validity perspective, and propose the use of the classic construct validity framework to evaluate the quality of text embeddings. First, we describe how this framework can be adapted to the opaque and high-dimensional nature of text embeddings. Second, we apply our adapted framework to an example where we compare the validity of survey question representation across text embedding models.

Keywords

Computational social science, Content validity, Convergent validity, Discriminant validity, Measurement validity, Predictive validity, Sentence embeddings, Survey methodology, Survey questions, Word embeddings, Modelling and Simulation, Computer Science Applications, Computational Mathematics

Citation

Fang, Q, Nguyen, D & Oberski, D L 2022, 'Evaluating the construct validity of text embeddings with application to survey questions', EPJ Data Science, vol. 11, no. 1, 39, pp. 1-31. https://doi.org/10.1140/epjds/s13688-022-00353-7