Parametrized Stochastic Grammars for RNA Secondary Structure Prediction
Robert S. Maier
Abstract
We propose a two-level stochastic context-free grammar (SCFG) architecture for parametrized stochastic modeling of a family of RNA sequences, including their secondary structure. A stochastic model of this type can be used for maximum a posteriori estimation of the secondary structure of any new sequence in the family. The proposed SCFG architecture models RNA subsequences comprising paired bases as stochastically weighted Dyck-language words, i.e., as weighted balanced-parenthesis expressions. The length of each run of unpaired bases, forming a loop or a bulge, is taken to have a phase-type distribution: that of the hitting time in a finite-state Markov chain. Without loss of generality, each such Markov chain can be taken to have a bounded complexity. The scheme yields an overall family SCFG with a manageable number of parameters.
Create a lesson
Related papers
SaltyMeta: a curated benchmark and protein language model-informed web tool for salty peptide prediction
Wanchao Chen, Wen Li, Yanan He et al.
Exploring Optimal Parameters for Ligand-Based Virtual Screening in Early Drug Discovery
Temitope Sobodu, Victor Chibuzor Johnson, Ryan Kern et al.
Synthesizing State-of-the-Art Structure Predictions from Soup of Co-folding Models
Hyosoon Jang, Taewon Kim, Sungsoo Ahn
Sequence-Informed Geometric Evaluation of RNA 3D Structures
Andrea Zerio, Yighua Yao, Alessandro Micheli et al.
Multi-ligand simultaneous docking of Carica papaya leaf phytochemicals, Carpaine and Rutin, reveals multi-mechanism inhibition of cancer proteins BCL-2 and WWP1
Merla Sudha, Asmita Saha, Belaguppa Manjunath Ashwin Desai et al.
Predicting directional flexibility in proteins
Vsevolod Viliuga, Leif Seute, Matteo Tadiello et al.