Non-extensive Trends in the Size Distribution of Coding and Non-coding DNA Sequences in the Human Genome
Th. Oikonomou, A. Provata
Abstract
We study the primary DNA structure of four of the most completely sequenced human chromosomes (including chromosome 19 which is the most dense in coding), using Non-extensive Statistics. We show that the exponents governing the decay of the coding size distributions vary between 5.2 r 5.7 for the short scales and 1.45 q 1.50 for the large scales. On the contrary, the exponents governing the decay of the non-coding size distributions in these four chromosomes, take the values 2.4 r 3.2 for the short scales and 1.50 q 1.72 for the large scales. This quantitative difference, in particular in the tail exponent q, indicates that the non-coding (coding) size distributions have long (short) range correlations. This non-trivial difference in the DNA statistics is attributed to the non-conservative (conservative) evolution dynamics acting on the non-coding (coding) DNA sequences.
Create a lesson
Related papers
Large Language Model Agents for Evidence Based Genetic Disease Severity Classification
Tohid Ghasemnejad, Ahmadreza Argha, Mark Grosser et al.
STUART: Sequence Triage and qUAntification of Read Transcripts for Rapid Ionizing Radiation Exposure Assessment
Tomasz Strzoda, Lourdes Cruz-Garcia, Mustafa Najim et al.
PlainMap: a lightweight, restartable mapping pipeline for ancient and modern DNA
Michael V. Westbury
Democratizing Clinical Tumor Whole Genome Sequencing: 18-hour End-to-end Analysis via Trillion-parameter Large Language Models Locally Deployed on Consumer-grade Hardware
Rui Xiao, Yili Xu
Structure is not mechanism: high-gain gated-FFN rows across text and genomic foundation models
Alexandros Tzanakakis, Aris Karatzikos, Ilias Georgakopoulos-Soares
RAGCell: Retrieval-Augmented Generation as Supervision for Versatile Single-cell Analysis
Tianyu Liu, Fan Zhang, Jiayuan Chen et al.