Statistical linguistic study of DNA sequences
K. L. Ng, S. P. Li
Abstract
A new family of compound Poisson distribution functions from statistical linguistic is used to study the n-tuples and nucleotide composition features of DNA sequences. The relative frequency distribution of the 6-tuples and 7- tuples occurrence studies suggest that most of the DNA sequences follow the general shape of the compound Poisson distribution. It is also noted that the χ-square test indicated that some of the sequences follow this distribution with a reasonable level of goodness of fit. The compositional segmentation study fits quite well using this new family of distribution functions. Furthermore, the absolute values of the relative frequency come out naturally from the linguistic model without ambiguity. It is suggesting that DNA sequences are not random sequences and they could possibly have subsequence structures.
Create a lesson
Related papers
Global Minima of the Thomson Problem in a Disk: A Molecular Dynamics Approach with Fixed Border Charges
Georgiy K. Lavrov, Eduard G. Nikonov
Martingale theory for heat and phase-space contraction in heterogeneous diffusions
Jing Qin, Nariya Uchida, Édgar Roldán
Formal Fluctuation-Response Relations for Non-Stationary Systems: The Dynamic Conjugate Variable
Igor M. Sokolov
Khinchin's ergodicity and typicality in statistical mechanics
Dario Lucente, Marco Baldovin, Giacomo Gradenigo et al.
Universal 1/f Noise in the Power Spectra of Energy Time-series in Solvated DNA Dynamics
Harsh Sahu, Deepika Sardana, Pramod Kumar et al.
Landau diamagnetism and the de Haas-van Alphen effect from a single geometric construction
Sung-Hoon Lee