Streaming algorithms for robust max-min diversification
Andrea Pietracaprina, Geppino Pucci, Stefano Zanon
Abstract
Given a set of n points X in a metric space and an integer k, max-min diversification aims to select k points of X maximizing their minimum pairwise distance. This objective function is however highly vulnerable to noisy points. In[Amagata, AAAI23], a robust formulation is proposed which addresses this vulnerability by excluding solutions containing any of z outliers, defined as the z points in X with the largest nearest-neighbor distances. That paper also presents a coreset-based streaming algorithm for the new formulation, based on a suitable inlier-outlier separation assumption. However, we identify three shortcomings in the algorithm by [Amagata, AAAI23]: its coreset construction requires an offline computation over X, which needs memory linear in n, in stark contrast with the typical goals of stream processing; the one-pass procedure used to extract the solution from the coreset may return fewer than k points (hence, an unfeasible solution) because it permanently discards points too far from the current solution; and its outlier-exclusion guarantee is only probabilistic and weakens as the coreset size shrinks. In contrast, we present a deterministic coreset-based algorithm that, under a natural inlier-outlier separation assumption (similar to the one used in [Amagata, AAAI23]), returns exactly k inliers which are a (2+)-approximate solution, for any >0, thus only above the best polynomial-time sequential approximation, even without outliers. Its one-pass streaming implementation adapts obliviously to the dataset's doubling dimension D and, for wide ranges of k, z, , and D, it uses memory independent of n. For sufficiently long streams, its amortized update time is proportional to the coreset size, thus also independent of n.
Create a lesson
Related papers
TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fine-Tuning
Jichao Jiang, Cristian McGee, El Houcine Bergou et al.
FERPO: Forward Entropy-Regularized Policy Optimization
Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv
Cost-augmented Schrödinger bridges on graphs are exactly solvable: a Feynman-Kac tilt replaces learned control
Akshay Balsubramani
The Missing Primitive: Diagnosing and Repairing Mathematical Reasoning in Large Language Models
Shuo Xing, Zilin Dai, Chengyuan Qian et al.
Trust the Direction, Search the Step: Zero-and-First-Order Methods for LLM Fine-Tuning
Cristian McGee, El Houcine Bergou, Aritra Dutta
Generative modeling of intrinsically disordered protein regions by reinforcing sparse autoencoder features
Jason X. Liu, Sebastian Ibarraran, Frank Hu et al.