reGenotyper: Detecting mislabeled samples in genetic data

Publication date

2017

Authors

Zych, Konrad
Snoek, BastenISNI 0000000419527486
Elvin, Mark
Rodriguez, Miriam
Van der Velde, K Joeri
Arends, Danny
Westra, Harm-Jan
Swertz, Morris A
Poulin, Gino
Kammenga, Jan E

Editors

Advisors

Supervisors

Document Type

Article
Open Access logo

License

Abstract

In high-throughput molecular profiling studies, genotype labels can be wrongly assigned at various experimental steps; the resulting mislabeled samples seriously reduce the power to detect the genetic basis of phenotypic variation. We have developed an approach to detect potential mislabeling, recover the "ideal" genotype and identify "best-matched" labels for mislabeled samples. On average, we identified 4% of samples as mislabeled in eight published datasets, highlighting the necessity of applying a "data cleaning" step before standard data analysis.

Keywords

Citation

Zych, K, Snoek, B L, Elvin, M, Rodriguez, M, Van der Velde, K J, Arends, D, Westra, H-J, Swertz, M A, Poulin, G, Kammenga, J E, Breitling, R, Jansen, R C & Li, Y 2017, 'reGenotyper : Detecting mislabeled samples in genetic data', PLoS One, vol. 12, no. 2, e0171324. https://doi.org/10.1371/journal.pone.0171324