REVIEW 5 cited by
A Survey on Locality Sensitive Hashing Algorithms and their Applications
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
A Survey on Locality Sensitive Hashing Algorithms and their Applications
read the original abstract
Finding nearest neighbors in high-dimensional spaces is a fundamental operation in many diverse application domains. Locality Sensitive Hashing (LSH) is one of the most popular techniques for finding approximate nearest neighbor searches in high-dimensional spaces. The main benefits of LSH are its sub-linear query performance and theoretical guarantees on the query accuracy. In this survey paper, we provide a review of state-of-the-art LSH and Distributed LSH techniques. Most importantly, unlike any other prior survey, we present how Locality Sensitive Hashing is utilized in different application domains.
Forward citations
Cited by 5 Pith papers
-
ASH: Asymmetric Scalar Hashing With Learned Dimensionality Reduction for High-Fidelity Vector Quantization
ASH achieves state-of-the-art ANN recall and speed across compression levels by learning an orthonormal projection for dimensionality reduction followed by scalar quantization in an asymmetric encoder-decoder setup.
-
H3D: Benchmarking Unsupervised Text Hashing for Fine-Grained Document Deduplication
Lexical non-learning hashes match near-duplicates well, while BGE-based quantized embeddings better preserve rewritten scientific similarity, under a shared ranking protocol on CSFCube and RELISH.
-
When More Cores Hurts: The Vector Database Scaling Paradox in HPC
Large-scale HPC evaluation of Qdrant, Milvus, and Weaviate reveals that workload patterns limit scaling and extra cores can reduce throughput, exposing a cloud-to-HPC design mismatch.
-
RcLLM: Accelerating Generative Recommendation via Beyond-Prefix KV Caching
RcLLM accelerates generative recommendation inference by 1.31x-9.51x in TTFT through beyond-prefix KV caching, replicated user caches, sharded item caches, affinity scheduling, and selective attention with negligible ...
-
Data Leakage Detection and De-duplication in Large Scale Geospatial Image Datasets
The AICrowd dataset has 90% training duplicates and 93% validation-to-training leakage; a perceptual hashing pipeline detects and mitigates these issues.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.