Information based clustering
Noam Slonim, Gurinder Singh Atwal, Gasper Tkacik, William Bialek
Abstract
In an age of increasingly large data sets, investigators in many different disciplines have turned to clustering as a tool for data analysis and exploration. Existing clustering methods, however, typically depend on several nontrivial assumptions about the structure of data. Here we reformulate the clustering problem from an information theoretic perspective which avoids many of these assumptions. In particular, our formulation obviates the need for defining a cluster "prototype", does not require an a priori similarity metric, is invariant to changes in the representation of the data, and naturally captures non-linear relations. We apply this approach to different domains and find that it consistently produces clusters that are more coherent than those extracted by existing algorithms. Finally, our approach provides a way of clustering based on collective notions of similarity rather than the traditional pairwise measures.
Create a lesson
Related papers
Surf2Volume: a workflow for converting CIFTI parcellations to NIfTI volume space
Shuguang Yang, Ziyi Wang, Yujing Shen et al.
DINIRS: Digital Twin for Individualized Treatment Effects of Non-Invasive Respiratory Support Strategies
Md Fantacher Islam, Jarrod Mosier, Vignesh Subbian
RegimeFormer: A Large Protein Model of Global Perturbation Regimes
Siyuan Ma, Yi Chai, Yi Wu et al.
Interpreting Latent Protein Language Model Features with Geometric Annotations
Siddharth Setlur, Djordje Mihajlovic, Darrick Lee
PathoMIC: A Benchmark for Cross-Species Antimicrobial Peptide Activity Prediction
Yeqing Lu, Xiaoyan Zhao, Fuli Feng
Multimodal risk trajectories reveal heterogeneous paths to dementia
Zhiqi Lee, Haowen Li, Tao Liu et al.