On the Complexity of Several Haplotyping Problems
Rudi Cilibrasi, Leo van Iersel, Steven Kelk, John Tromp
Abstract
In this paper we present a collection of results pertaining to haplotyping. The first set of results concerns the combinatorial problem of reconstructing haplotypes from incomplete and/or imperfectly sequenced haplotype data. More specifically, we show that an interesting, restricted case of Minimum Error Correction (MEC) is NP-hard, point out problems in earlier claims about a related problem, and present a polynomial-time algorithm for the ungapped case of Longest Haplotype Reconstruction (LHR). Secondly, we present a polynomial time algorithm for the problem of resolving genotype data using as few haplotypes as possible (the Pure Parsimony Haplotyping Problem, PPH) where each genotype has at most two ambiguous positions, thus solving an open problem posed by Lancia et al in "Haplotyping Populations by Pure Parsimony: Complexity of Exact and Approximation Algorithms."
Create a lesson
Related papers
Large Language Model Agents for Evidence Based Genetic Disease Severity Classification
Tohid Ghasemnejad, Ahmadreza Argha, Mark Grosser et al.
STUART: Sequence Triage and qUAntification of Read Transcripts for Rapid Ionizing Radiation Exposure Assessment
Tomasz Strzoda, Lourdes Cruz-Garcia, Mustafa Najim et al.
PlainMap: a lightweight, restartable mapping pipeline for ancient and modern DNA
Michael V. Westbury
Democratizing Clinical Tumor Whole Genome Sequencing: 18-hour End-to-end Analysis via Trillion-parameter Large Language Models Locally Deployed on Consumer-grade Hardware
Rui Xiao, Yili Xu
Structure is not mechanism: high-gain gated-FFN rows across text and genomic foundation models
Alexandros Tzanakakis, Aris Karatzikos, Ilias Georgakopoulos-Soares
RAGCell: Retrieval-Augmented Generation as Supervision for Versatile Single-cell Analysis
Tianyu Liu, Fan Zhang, Jiayuan Chen et al.