Are You Learning Biological Signal or Shortcuts? Auditing and Mitigating Bias in Protein-Protein Interaction Datasets
Judith Bernett, Anton Spannagl, Joel Ås, Markus List, David B. Blumenthal
Abstract
Protein-protein interaction (PPI) databases do not faithfully reflect biological realities. Instead, they are influenced by study and technical biases that distort certain protein and interaction attributes. Machine learning models can exploit these as learning shortcuts if the negative dataset is not constructed with care. So far, the shortcuts introduced during PPI dataset construction have only been examined in isolation. Here, we systematically characterize both reported and, to our knowledge, previously unreported biases in PPI datasets that lead machine learning models to learn shortcuts instead of biological signal. We analyze HIPPIE, IntAct, and STRING, dedicated PPI databases, as well as two datasets derived from 3D-structural information in the Protein Data Bank (PDB). We show that random data splitting introduces strong topological shortcuts. When train-test protein overlap is removed, the resulting datasets still retain usable shortcuts stemming from self-interactions, taxonomic identity, and functional relatedness, whose prevalence interestingly depends on the data source. We further show that sampling negatives from a set of high-confidence non-interactors, an intuitively appealing choice, can amplify the shortcut stemming from functional relatedness. To detect and mitigate these biases, we provide an open Nextflow pipeline that combines similarity-aware, data-loss-minimizing dataset splitting with bias-minimizing negative sampling, both formulated as integer linear programs. Its key concept of quantifying biases to minimize them through optimization-based negative sampling can, in principle, be extended to any machine learning problem where the pool of negative candidates is much larger than the positives and is thus of interest also beyond PPI prediction.
Create a lesson
Related papers
Local energetic coupling enhances the expressivity of chemical computation
Marco Tuccio, Jason W. Rocks, Joshua E. Goldford
Thermodynamic and Statistical Signatures of Modality Changes in Concentration Distributions Driven by Stochastic Switching Between Two Activity States
Aindrila Deb, Pintu Patra
Orchestra: Corroboration-Based Regulatory Candidate Discovery via Composed Bioinformatics MCP Agents
Jose A. Bird
Systematic pathway comparison on the powerset of rule-based biochemical systems
Anne-Susann Abel, Sissel Banke, Erika M. Herrera Machado et al.
Uncovering Cellular Resolution in scRNAseq via Unbiased Cell and Gene Network Analysis
Olga lanzetta, Luisa Cutillo, Bailey Andrew et al.
DigiPhen: a new paradigm for building predictive models of biological systems
H. Steven Wiley, Angela Cintolesi, Niaz Bahar Chowdhury et al.