PCA and K-Means decipher genome
A. N. Gorban, A. Y. Zinovyev
Abstract
In this paper, we aim to give a tutorial for undergraduate students studying statistical methods and/or bioinformatics. The students will learn how data visualization can help in genomic sequence analysis. Students start with a fragment of genetic text of a bacterial genome and analyze its structure. By means of principal component analysis they ``discover'' that the information in the genome is encoded by non-overlapping triplets. Next, they learn how to find gene positions. This exercise on PCA and K-Means clustering enables active study of the basic bioinformatics notions. Appendix 1 contains program listings that go along with this exercise. Appendix 2 includes 2D PCA plots of triplet usage in moving frame for a series of bacterial genomes from GC-poor to GC-rich ones. Animated 3D PCA plots are attached as separate gif files. Topology (cluster structure) and geometry (mutual positions of clusters) of these plots depends clearly on GC-content.
Create a lesson
Related papers
Automatic denoising and differentiation based on Savitzky-Golay filtering and Homogeneous Differentiators for attractor reconstruction via differential embedding
Uros Sutulovic, Daniele Proverbio, Rami Katz et al.
FlowLOT: Linearized Optimal Transport for Flow Cytometry Analysis
Naqib Sad Pathan, Mohammad Shifat-E-Rabbi, Kristofor E. Pas et al.
GIA: Germline-Informed Aging with AlphaGenome Finds Genetically Regulated CpGs
Sean Lim
Decoding Extrahepatic Targeting of Lipid Nanoparticles with Interpretable Machine Learning
Asal Mehradfar, Mohammad Shahab Sepehri, Owen Antholine et al.
GPCR Ligand Bioactivity Prediction with Physics-Informed Dual-State Query Learning
Shuo Zhang, Huifeng Zhang, Rongqi Hong et al.
Optical microelectrode arrays for differential readout of electrical and mechanical signals in cardiac cells
Alessandro Leronni, Rosalia Moreddu