Efficient Auto-Interpretability of AI Models in Biology
Piotr Jedryszek, Oliver M. Crook
Abstract
Sparse autoencoders (SAEs), and other interpretability methods could turn AI models in Biology and other fields into engines of scientific discovery by explaining the superhuman capabilities of those models. However, a latent is only useful if we know three things: whether it is coherent, whether it can be described, and whether that description has predictive power. These questions are routinely conflated. We assemble them into a single pipeline and report the practical innovations each stage required. First, cross-seed dictionary stability prioritises which latents are worth spending resources to investigate. Second, an intruder-detection task asks whether a latents activating examples share a recognizable pattern. Third, a separate pass proposes a candidate biological description which we convert into falsifiable predictions which can be tested in silico. Deployed on the Boltz-1 Pairformer trunk, stability prioritisation finds interpretable latents using about 4.4 times fewer latent evaluations each, and at 5.2 times lower measured cost, while recovering over half of them, and the external check shows the surfaced motifs are significantly enriched for their claimed annotations. The results also suggest a possible tension: the cross- seed stability might be selecting for some types of features, like structure-related ones, much more than others, such as function-related features.
Create a lesson
Related papers
Science sandboxes measure the scientific capability of AI agents
Arya S. Rao, Rodrigo I. Castro, Sager J. Gosai et al.
CryoAnomaly: Few-Shot Cryo-EM Particle Picking via Anomaly-Guided Hard Negative Suppression
Riku Itsuji, Rintaro Otsubo, Ryo Fujii et al.
FoldKit: A Python library for efficient storage and retrieval of co-folding predictions
Jonathan A. Levine, Melissa Pathil, Samuel Nitz et al.
CIR-DDG: backbone-agnostic residual correction of antibody-antigen affinity changes with explicit cross-chain geometry
Weilun Yu, Zhiheng Zou, Yonggui Huang et al.
Surf2Volume: a workflow for converting CIFTI parcellations to NIfTI volume space
Shuguang Yang, Ziyi Wang, Yujing Shen et al.
DINIRS: Digital Twin for Individualized Treatment Effects of Non-Invasive Respiratory Support Strategies
Md Fantacher Islam, Jarrod Mosier, Vignesh Subbian