Skip to content

misi: a Metric Inverted Sample Index

Edgar Chavez

cs.IRarXiv:2608.27422

Abstract

We present misi, an inverted index for approximate nearest-neighbor search over general metric spaces whose vocabulary is a random sample of the database, of size proportional to n. Each object is represented by its kb nearest sample points, found by a pluggable inner index over the sample; queries are answered by an idf-weighted shared-neighbor vote followed by exact verification of C candidates. The construction generalizes the NAPP index from a constant number of pivots to a linear-size vocabulary, which keeps posting lists at constant expected length ρ= kb/α as n grows and turns the index into a combinator: any high-recall index on αn points yields an index on n points, for any metric. A probabilistic model gives a recall guarantee -- kb logarithmic in n over the overlap gap suffices, with a verification budget the index itself estimates -- and a matching limit: the vote cannot resolve overlap differences below order 1/kb. The design's strengths are structural: construction is n independent searches -- embarrassingly parallel, deterministic, 5,250 s for 108 vectors on 64 cores, 3.7× faster than a matched-recall graph build -- it streams under an enforced 3 GiB cap, and the portable artifact serves 108 vectors from NVMe within an enforced 8 GB budget, below the working floor of the SSD-graph baseline. Its cost is query-time work: saturated graph baselines answer 6-16× faster in RAM, and the verification budget for 0.99 recall grows as n0.30. All results carry seeds, saturation sweeps and full configurations, are generated from run manifests, and include measured negative results. The intended applications weight construction cost, determinism, memory footprint, or black-box metrics over peak throughput: frequently rebuilt corpora, batch similarity workloads, constrained-memory serving.

Create a lesson