misi: a Metric Inverted Sample Index
Edgar Chavez
Abstract
We present misi, an inverted index for approximate nearest-neighbor search over general metric spaces whose vocabulary is a random sample of the database, of size proportional to n. Each object is represented by its kb nearest sample points, found by a pluggable inner index over the sample; queries are answered by an idf-weighted shared-neighbor vote followed by exact verification of C candidates. The construction generalizes the NAPP index from a constant number of pivots to a linear-size vocabulary, which keeps posting lists at constant expected length ρ= kb/α as n grows and turns the index into a combinator: any high-recall index on αn points yields an index on n points, for any metric. A probabilistic model gives a recall guarantee -- kb logarithmic in n over the overlap gap suffices, with a verification budget the index itself estimates -- and a matching limit: the vote cannot resolve overlap differences below order 1/kb. The design's strengths are structural: construction is n independent searches -- embarrassingly parallel, deterministic, 5,250 s for 108 vectors on 64 cores, 3.7× faster than a matched-recall graph build -- it streams under an enforced 3 GiB cap, and the portable artifact serves 108 vectors from NVMe within an enforced 8 GB budget, below the working floor of the SSD-graph baseline. Its cost is query-time work: saturated graph baselines answer 6-16× faster in RAM, and the verification budget for 0.99 recall grows as n0.30. All results carry seeds, saturation sweeps and full configurations, are generated from run manifests, and include measured negative results. The intended applications weight construction cost, determinism, memory footprint, or black-box metrics over peak throughput: frequently rebuilt corpora, batch similarity workloads, constrained-memory serving.
Create a lesson
Related papers
Scaling Graph Neural Networks for Friend Recommendation: Multi-Hash User Embeddings and Temporal Neighbor Sampling
Maksim Utushkin, Andrei Ovsiannikov, Alexander D'yakonov
Stageboost: Recommending Signals Based on Counterfactual Estimation
Darpan Singhal, Matan Mandelbrod, Tal Franji et al.
Astar: Learning to Propose Evolution Directions for Self-Evolving Industrial AI Systems
Jinxin Hu, Hao Deng, Haibo Xing et al.
ProRetrieval: Learning to Orchestrate Hybrid Search via Executable Program Synthesis
Chengsong You, Zhen Sun, Yunhai Hu et al.
Conversational Recommendation over Live E-Commerce Catalogues with Self-Refreshing Retrieval
Ante Kapetanovic, Tomislav Duricic, Dionizije Fa et al.
Topology-Masked Unified Backbone for Joint Feature Interaction and Multi-Domain Sequence Modeling
Zhihao Zhu, Dezheng Han, Jikang Xia et al.