Efficient estimation of the cardinality of large data sets
Philippe Chassaing, Lucas Gerin
Abstract
F.Giroire has recently proposed an algorithm which returns the approximate number of distincts elements in a large sequence of words, under strong constraints coming from the analysis of large data bases. His estimation is based on statistical properties of uniform random variables in [0,1]. In this note we propose an optimal estimation, using Kullback information and estimation theory.
Create a lesson
Related papers
Instance-Optimal Adaptive Location Estimation via Multiscale Mid-Summaries
Qiaosen Wang, Chao Gao
Robust Multi-Task Learning for Principal Component Analysis
Dali Liu, Haolei Weng
Principal component error in high-dimensional factor models
Alex Bernstein, Lisa R. Goldberg, Nicholas Gunther et al.
Approximation Theorems for High-Dimensional Canonical U-Statistics: Gaussian Chaos and Phase Transition
Leheng Cai, Qirui Hu
On the parametric and semiparametric Fisher information matrix for non-zero mean stationary spherical invariant random processes
Jean-Pierre Delmas, Habti Abeida, Stefano Fortunati
Inference for two-stage sampling in spatial surveys
Guillaume Chauvet, Olivier Bouriaud, Trinh H. K. Duong