Are knowledge graph embedding models biased, or is it the data that they are trained on?

Publication date

2021-10

Authors

Radstok, WesselISNI 0000000507296649
Chekol, Melisachew WudageISNI 0000000433166787
Schaefer, Mirko TobiasORCID 0000-0003-0212-7016ISNI 0000000356270811

Editors

Advisors

Supervisors

DOI

Document Type

Contribution to conference
Open Access logo

License

cc_by

Abstract

Recent studies on bias analysis of knowledge graph (KG) embedding models focus primarily on altering the models such that sensitive features are dealt with differently from other features. The underlying implication is that the models cause bias, or that it is their task to solve it. In this paper we argue that the problem is not caused by the models but by the data, and that it is the responsibility of the expert to ensure that the data is representative for the intended goal. To support this claim, we experiment with two different knowledge graphs and show that the bias is not only present in the models, but also in the data. Next, we show that by adding new samples to balance the distribution of facts with regards to specifc sensitive features, we can reduce the bias in the models.

Keywords

Citation

Radstok, W, Chekol, M & Schaefer, M 2021, 'Are knowledge graph embedding models biased, or is it the data that they are trained on?', Paper presented at Wikidata Workshop 2021 co-located with the 20th International Semantic Web Conference (ISWC 2021), 24/10/21 - 24/10/21. < http://ceur-ws.org/Vol-2982/ >, conference