Overlapping Probabilities of Top Ranking Gene Lists, Hypergeometric Distribution, and Stringency of Gene Selection Criterion
Wen Fury, Franak Batliwalla, Peter K. Gregersen, Wentian Li
Abstract
When the same set of genes appear in two top ranking gene lists in two different studies, it is often of interest to estimate the probability for this being a chance event. This overlapping probability is well known to follow the hypergeometric distribution. Usually, the lengths of top-ranking gene lists are assumed to be fixed, by using a pre-set criterion on, e.g., p-value for the t-test. We investigate how overlapping probability changes with the gene selection criterion, or simply, with the length of the top-ranking gene lists. It is concluded that overlapping probability is indeed a function of the gene list length, and its statistical significance should be quoted in the context of gene selection criterion.
Create a lesson
Related papers
Surf2Volume: a workflow for converting CIFTI parcellations to NIfTI volume space
Shuguang Yang, Ziyi Wang, Yujing Shen et al.
DINIRS: Digital Twin for Individualized Treatment Effects of Non-Invasive Respiratory Support Strategies
Md Fantacher Islam, Jarrod Mosier, Vignesh Subbian
RegimeFormer: A Large Protein Model of Global Perturbation Regimes
Siyuan Ma, Yi Chai, Yi Wu et al.
Interpreting Latent Protein Language Model Features with Geometric Annotations
Siddharth Setlur, Djordje Mihajlovic, Darrick Lee
PathoMIC: A Benchmark for Cross-Species Antimicrobial Peptide Activity Prediction
Yeqing Lu, Xiaoyan Zhao, Fuli Feng
Multimodal risk trajectories reveal heterogeneous paths to dementia
Zhiqi Lee, Haowen Li, Tao Liu et al.