Pith. sign in

REVIEW 4 cited by

Spreading vectors for similarity search

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1806.03198 v3 pith:3AMVYPD5 submitted 2018-06-08 stat.ML cs.LG

classification stat.MLcs.LG
keywords dataquantizermethodsneuralproposequantizationquantizersstep
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Discretizing multi-dimensional data distributions is a fundamental step of modern indexing methods. State-of-the-art techniques learn parameters of quantizers on training data for optimal performance, thus adapting quantizers to the data. In this work, we propose to reverse this paradigm and adapt the data to the quantizer: we train a neural net which last layer forms a fixed parameter-free quantizer, such as pre-defined points of a hyper-sphere. As a proxy objective, we design and train a neural network that favors uniformity in the spherical latent space, while preserving the neighborhood structure after the mapping. We propose a new regularizer derived from the Kozachenko--Leonenko differential entropy estimator to enforce uniformity and combine it with a locality-aware triplet loss. Experiments show that our end-to-end approach outperforms most learned quantization methods, and is competitive with the state of the art on widely adopted benchmarks. Furthermore, we show that training without the quantization step results in almost no difference in accuracy, but yields a generic catalyzer that can be applied with any subsequent quantizer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 8 citations worldwide. Full citation record

  1. SleepMaMi: A Universal Sleep Foundation Model for Integrating Macro- and Micro-structures

    cs.AI 2026-02 conditional novelty 6.0 of 10

    SleepMaMi, a dual-encoder sleep foundation model pretrained on 158K hours of PSG, matches or beats existing sleep foundation models on staging, apnea segmentation, and disease prediction.

  2. Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A foundation model of wearable behavioral data outperforms simple baselines and complements a PPG sensor model across 57 health detection tasks.

  3. Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning

    cs.CV 2025-06 reject novelty 6.0 of 10

    AsymDSD unifies latent masked point modeling and cross-view invariance self-distillation to learn 3D representations, reporting 90.53% on ScanObjectNN and 93.72% with 930k-shape pretraining.

  4. PiPViT: Patch-based Visual Interpretable Prototypes for Retinal Image Analysis

    cs.CV 2025-06 conditional novelty 4.0 of 10

    PiPViT combines vision transformers and prototype learning to classify retinal OCT scans while showing the spatial extent of the biomarker that drove the decision.

Pith tools