Pith. sign in

REVIEW 3 cited by

SOAR: Improved Indexing for Approximate Nearest Neighbor Search

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.00774 v1 pith:4YEJJAL5 submitted 2024-03-31 cs.LG

classification cs.LG
keywords searchsoarindexingnearestneighborrepresentationsapproximatedata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces SOAR: Spilling with Orthogonality-Amplified Residuals, a novel data indexing technique for approximate nearest neighbor (ANN) search. SOAR extends upon previous approaches to ANN search, such as spill trees, that utilize multiple redundant representations while partitioning the data to reduce the probability of missing a nearest neighbor during search. Rather than training and computing these redundant representations independently, however, SOAR uses an orthogonality-amplified residual loss, which optimizes each representation to compensate for cases where other representations perform poorly. This drastically improves the overall index quality, resulting in state-of-the-art ANN benchmark performance while maintaining fast indexing times and low memory consumption.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CleANN: Efficient Full Dynamism in Graph-based Approximate Nearest Neighbor Search

    cs.DB 2025-07 conditional novelty 7.0 of 10

    CleANN combines workload-aware bridge building, on-the-fly neighborhood consolidation, and semi-lazy memory cleaning to keep graph-based ANNS recall near static-build levels under fully dynamic concurrent workloads.

  2. Vectorizing the Trie: Efficient Constrained Decoding for LLM-based Generative Retrieval on Accelerators

    cs.IR 2026-02 conditional novelty 6.0 of 10

    Constrained decoding for generative retrieval can be made accelerator-friendly by flattening the trie of valid items into a CSR sparse matrix and doing branch-free vectorized lookups.

  3. kANNolo: Sweet and Smooth Approximate k-Nearest Neighbors Search

    cs.IR 2025-01 conditional novelty 5.0 of 10

    kANNolo, a modular Rust ANN library built on HNSW and product quantization, achieves state-of-the-art speed-accuracy trade-offs on dense and sparse benchmarks.

Pith tools