Journal Article

·2012

Statistical bias and variance of gene selection and cross validation methods: A case study on hypertension prediction

Zeliha Görmez , Olcay Kurşun , Ahmet Sertbaş , Nizamettin Aydın YTU , Hüseyin Şeker

Abstract

In exploratory association studies of genes with certain diseases, a single or a small number of genes (features) related with the diseases are selected 1 among many thousands investigated. We investigate the statistical bias and variance of simple yet common (correlation and mutual information based) feature selection algorithms using well-known cross-validation methods (leave-one-out and k-fold) on a gene finding study for hypertension prediction. Our findings show that selected genes are different for different methods and different cross-validation runs for both single gene selection and gene subset selection.

Keywords

Cross-validation Selection (genetic algorithm) Feature selection Gene Variance (accounting) Computational biology Gene selection Computer science Correlation Selection bias Statistics Data mining Artificial intelligence Biology Genetics Mathematics Gene expression

Subject Areas

Gene expression and cancer classification ·Molecular Biology ·Life Sciences
Bioinformatics and Genomic Networks ·Molecular Biology ·Life Sciences
Genetic Associations and Epidemiology ·Genetics ·Life Sciences

Citations by Year

OpenAlex SDG Match

SDGs auto-classified by OpenAlex (score ≥ 0.4 shown).

Quality Education 54%