Exact Limits of Random Projections for Preserving Geometry: Distance Recovery, Nearest-Neighbor Rankings, and Covariance Shape in Gaussian Models
Piyush Sao
Abstract
The Johnson-Lindenstrauss (JL) lemma guarantees that a random projection of n points to m=O(-2 n) dimensions preserves pairwise squared distances within relative error with high probability, and this dimension order is asymptotically optimal. In high dimensions, however, distances concentrate around a baseline while key geometric information lies in much smaller fluctuations. We show that the JL bound can therefore be uninformative about retained geometry: an independent Gaussian replacement map can satisfy it even though the replacement cloud is independent of the original data. We then ask how well any decoder can recover a feature f(D) of a squared distance D from a linear sketch. Under squared-error loss, the optimal decoder is conditional expectation, so recovery defines a linear operator whose singular values quantify feature recovery. For isotropic Gaussian data (Σ=σ2 Id), we diagonalize this operator in closed form. For fixed k with m,d-m∞, its kth singular value satisfies k≈(m/ d)k/2. This yields three sharp consequences. A rank-m sketch retains at most an m/d fraction of the variance of any feature of one squared distance. If m∞ and m/d0, the expected Kendall correlation is 2πm/d(1+o(1)); for fixed q, nearest- neighbor agreement tends to 1/q. Yet one projection can satisfy the JL bound while mean Kendall correlation vanishes when n m d. After removing scale, Haar-averaged retained covariance-shape information is (m/d)2. Thus JL distance preservation does not quantify the geometry available for comparison or inference.
Create a lesson
Related papers
A Common Measure of Communication for Speech Brain-Computer Interfaces
Dulhan Jayalath, Benjamin Ballyk, Oiwi Parker Jones
Graph Machine: Towards Better Pretraining via Edges
Lintai Hou
The Implications of Linguistic Illegibility for LLM Security
James Mickens
Post-Training Language Models for Gold-Medal Performance in Coding Competitions
Aleksander Ficek, Sean Narenthiran, Mehrzad Samadi et al.
UE5M3 FP4 Block Scaling for Stable Language Model Pretraining
Robert Hu, Carlo Luschi, Paul Balanca
Cliff: Learning Process Rewards from the First Mistake
Peixuan Han, Runhui Wang, Ketan Ramaneti et al.