Abstract
In exploratory association studies of genes with certain diseases, a single or a small number of genes (features) related with the diseases are selected 1 among many thousands investigated. We investigate the statistical bias and variance of simple yet common (correlation and mutual information based) feature selection algorithms using well-known cross-validation methods (leave-one-out and k-fold) on a gene finding study for hypertension prediction. Our findings show that selected genes are different for different methods and different cross-validation runs for both single gene selection and gene subset selection.
Keywords
Subject Areas
Citations by Year
OpenAlex SDG Match
SDGs auto-classified by OpenAlex (score ≥ 0.4 shown).