Information Capacity of Biological Macromoleculae Reloaded
Michael G. Sadovsky
Abstract
Information capacity of a symbol sequence is a measure of the unexpectedness of a continuation of given string of symbols. Continuation of a string is determined through the maximum entropy of the reconstructed frequency dictionary; the capacity, in turn, is determined through the calculation of mutual entropy of a real frequency dictionary of a sequence with respect to the reconstructed one. The capacity does not depend on the length of strings in a dictionary. The capacity calculated for various genomes exhibits a multi-minima pattern reflecting an order observed within a sequence.
Create a lesson
Related papers
Large Language Model Agents for Evidence Based Genetic Disease Severity Classification
Tohid Ghasemnejad, Ahmadreza Argha, Mark Grosser et al.
STUART: Sequence Triage and qUAntification of Read Transcripts for Rapid Ionizing Radiation Exposure Assessment
Tomasz Strzoda, Lourdes Cruz-Garcia, Mustafa Najim et al.
PlainMap: a lightweight, restartable mapping pipeline for ancient and modern DNA
Michael V. Westbury
Democratizing Clinical Tumor Whole Genome Sequencing: 18-hour End-to-end Analysis via Trillion-parameter Large Language Models Locally Deployed on Consumer-grade Hardware
Rui Xiao, Yili Xu
Structure is not mechanism: high-gain gated-FFN rows across text and genomic foundation models
Alexandros Tzanakakis, Aris Karatzikos, Ilias Georgakopoulos-Soares
RAGCell: Retrieval-Augmented Generation as Supervision for Versatile Single-cell Analysis
Tianyu Liu, Fan Zhang, Jiayuan Chen et al.