Subcellularly Resolved Single-Cell Embedding Learning with Transcriptomic data, Protein Structure and Localization Information
Zhen Zhou, Jiachen Li, Yuan Liu, Xiaoyong Pan, Hong-Bin Shen
Abstract
Existing cell embedding methods predominantly rely on transcriptomic or proteomic measurements and represent each cell as a holistic entity, thereby overlooking the subcellular localization of individual molecules. Moreover, they rarely incorporate protein structural information, despite its fundamental role in determining molecular interactions and functions. In this work, we propose a multimodal framework for learning subcellularly resolved cell embeddings by jointly leveraging RNA expression profiles, protein sequence representations, and protein structural information. Specifically, we employ a cross-attention architecture to integrate transcriptomic, sequence, and structural modalities and model their interactions within distinct subcellular compartments. The resulting embeddings represent each cell through its fine-grained subcellular organization, capturing both molecular expression patterns and the functional properties of the associated proteins. By learning cell representations at subcellular resolution, our framework preserves spatially organized biological information while integrating complementary signals across multiple molecular levels. To the best of our knowledge, this is the first framework that produces subcellularly resolved cell embeddings by jointly incorporating transcriptomic information, protein sequence representations, and protein structural knowledge within a unified cross-modal learning paradigm.
Create a lesson
Related papers
PopPert: Population-level Joint-Distribution Modeling for Single-Cell Perturbation Prediction
Handong Wang, Jiaxin Qi, Haochen Feng et al.
Confounder-Aware Feature Correction for Single-Cell Batch Integration
Calvin McCarter
Learning Interpretable Tumor Microenvironment Representations by Fitting Pan-Cancer Cell State-Niche Correlation
Xiao Xiao, Jiashu He, Shiyang Zhang et al.
Optimizing RNA yield using deep neural networks coupled to massively parallel screening
Dinghai Zheng, Justin Hong, Jun Wang et al.
A Conditional Structure-Aware Generative Transformer for Multi-Objective Design of m1Ψ-Modified RNA 5' UTRs
Narges Zarnaghinaghsh, Ahmadreza Mofayezi, Byung-Jun Yoon
mLS-GKM: Efficient Multi-class Regulatory Sequence Classification with Gapped k-mer SVMs
Kieran Howard, Nathan Harmston