Ultrametric embedding: application to data fingerprinting and to fast data clustering
Fionn Murtagh
Abstract
We begin with pervasive ultrametricity due to high dimensionality and/or spatial sparsity. How extent or degree of ultrametricity can be quantified leads us to the discussion of varied practical cases when ultrametricity can be partially or locally present in data. We show how the ultrametricity can be assessed in text or document collections, and in time series signals. An aspect of importance here is that to draw benefit from this perspective the data may need to be recoded. Such data recoding can also be powerful in proximity searching, as we will show, where the data is embedded globally and not locally in an ultrametric space.
Create a lesson
Related papers
Conformal Prediction Through the Lens of Hypothesis Testing: Universality, Impossibility, and Optimality
Ryan J. Tibshirani, Rina Foygel Barber, Aaditya Ramdas
Connecting Riemannian Geometry and Statistical Inference for Correlation Matrices
Argyn Kuketayev
How far can symmetry help? Phase transitions and symmetry selection in sparse functional data analysis
Jocelyn Nembe
Posterior consistency for subdiffusion inverse problems
Haoyu Lu, Shaokang Zu, Junxiong Jia
Dimension comparison for Student's statistic under symmetric unimodality
Jacopo Lenzi
Statistical Properties of Nonparametric MLE under Laplace Noise
Yifei Xiong, Nianqiao Phyllis Ju, Vinayak Rao