REVIEW 7 cited by
Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Contrastive representation learning has been outstandingly successful in practice. In this work, we identify two key properties related to the contrastive loss: (1) alignment (closeness) of features from positive pairs, and (2) uniformity of the induced distribution of the (normalized) features on the hypersphere. We prove that, asymptotically, the contrastive loss optimizes these properties, and analyze their positive effects on downstream tasks. Empirically, we introduce an optimizable metric to quantify each property. Extensive experiments on standard vision and language datasets confirm the strong agreement between both metrics and downstream task performance. Remarkably, directly optimizing for these two metrics leads to representations with comparable or better performance at downstream tasks than contrastive learning. Project Page: https://tongzhouwang.info/hypersphere Code: https://github.com/SsnL/align_uniform , https://github.com/SsnL/moco_align_uniform
Forward citations
Cited by 7 Pith papers
-
Similarity search generalisation in contrastive learning with InfoNCE loss
InfoNCE risk with k negatives is O(1/k)-close to expected cross-entropy of softmax similarity search versus the positive generator, and a new continuity bound shows averaging stabilises generalisation as k grows.
-
Foundation Model-Driven Semantic Change Detection in Remote Sensing Imagery
PerASCD sets new state-of-the-art Sek scores on SECOND and LandsatSCD datasets by using a modular cascaded gated decoder on PerA foundation model features plus a new consistency loss.
-
A Geometry-Aware Efficient Algorithm for Compositional Entropic Risk Minimization
SCENT, a stochastic proximal mirror descent on the dual variable with an exponential Bregman divergence, optimizes compositional entropic risk at O(1/sqrt(T)) in the convex setting and matches or beats baselines on la...
-
The Hyperspherical Geometry of CLIP Latent Space: A Semantic Mixture Model
CLIP embeddings are modeled as a mixture of von Mises-Fisher distributions on the unit sphere, improving out-of-distribution detection and semantic decomposition over single-Gaussian baselines.
-
On the rankability of visual embeddings
Visual embeddings from CLIP and other vision encoders encode ordinal attributes along linear directions, recoverable from as few as two extreme reference images, without full supervision.
-
FedGraM: Defending Against Untargeted Attacks in Federated Learning via Embedding Gram Matrix
A server that keeps one example per class can detect and drop malicious federated-learning clients by measuring how separated their learned embeddings are, via the norm of a Gram matrix.
-
Skill Learning via Policy Diversity Yields Identifiable Representations for Reinforcement Learning
Contrastive Successor Features recover ground-truth RL states up to a linear map whenever the skill-conditioned transition differences follow a von Mises-Fisher distribution and policies are diverse.
Discussion (0). Continue with ORCID to comment.