Impact of pre-imputation SNP-filtering on genotype imputation results

Abstract:

BACKGROUND Imputation of partially missing or unobserved genotypes is an indispensable tool for SNP data analyses. However, research and understanding of the impact of initial SNP-data quality control on imputation results is still limited. In this paper, we aim to evaluate the effect of different strategies of pre-imputation quality filtering on the performance of the widely used imputation algorithms MaCH and IMPUTE. RESULTS We considered three scenarios: imputation of partially missing genotypes with usage of an external reference panel, without usage of an external reference panel, as well as imputation of completely un-typed SNPs using an external reference panel. We first created various datasets applying different SNP quality filters and masking certain percentages of randomly selected high-quality SNPs. We imputed these SNPs and compared the results between the different filtering scenarios by using established and newly proposed measures of imputation quality. While the established measures assess certainty of imputation results, our newly proposed measures focus on the agreement with true genotypes. These measures showed that pre-imputation SNP-filtering might be detrimental regarding imputation quality. Moreover, the strongest drivers of imputation quality were in general the burden of missingness and the number of SNPs used for imputation. We also found that using a reference panel always improves imputation quality of partially missing genotypes. MaCH performed slightly better than IMPUTE2 in most of our scenarios. Again, these results were more pronounced when using our newly defined measures of imputation quality. CONCLUSION Even a moderate filtering has a detrimental effect on the imputation quality. Therefore little or no SNP filtering prior to imputation appears to be the best strategy for imputing small to moderately sized datasets. Our results also showed that for these datasets, MaCH performs slightly better than IMPUTE2 in most scenarios at the cost of increased computing time.

DOI: 10.1186/s12863-014-0088-5

Projects: Genetical Statistics and Systems Biology

Publication type: Journal article

Journal: BMC genetics

Human Diseases: No Human Disease specified

Citation: BMC Genet 15(1),88

Date Published: 1st Dec 2014

Registered Mode: imported from a bibtex file

Authors: Nab Raj Roshyara, Holger Kirsten, Katrin Horn, Peter Ahnert, Markus Scholz

Help
help Submitter
Citation
Roshyara, N. R., Kirsten, H., Horn, K., Ahnert, P., & Scholz, M. (2014). Impact of pre-imputation SNP-filtering on genotype imputation results. In BMC Genetics (Vol. 15, Issue 1). Springer Science and Business Media LLC. https://doi.org/10.1186/s12863-014-0088-5
Activity

Views: 719

Created: 14th Sep 2020 at 13:35

Last updated: 7th Dec 2021 at 17:58

help Tags

This item has not yet been tagged.

help Attributions

None

Related items

Powered by
(v.1.13.0-master)
Copyright © 2008 - 2021 The University of Manchester and HITS gGmbH
Institute for Medical Informatics, Statistics and Epidemiology, University of Leipzig

By continuing to use this site you agree to the use of cookies