REVIEW 2 cited by
On the duality between contrastive and non-contrastive self-supervised learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recent approaches in self-supervised learning of image representations can be categorized into different families of methods and, in particular, can be divided into contrastive and non-contrastive approaches. While differences between the two families have been thoroughly discussed to motivate new approaches, we focus more on the theoretical similarities between them. By designing contrastive and covariance based non-contrastive criteria that can be related algebraically and shown to be equivalent under limited assumptions, we show how close those families can be. We further study popular methods and introduce variations of them, allowing us to relate this theoretical result to current practices and show the influence (or lack thereof) of design choices on downstream performance. Motivated by our equivalence result, we investigate the low performance of SimCLR and show how it can match VICReg's with careful hyperparameter tuning, improving significantly over known baselines. We also challenge the popular assumption that non-contrastive methods need large output dimensions. Our theoretical and quantitative results suggest that the numerical gaps between contrastive and non-contrastive methods in certain regimes can be closed given better network design choices and hyperparameter tuning. The evidence shows that unifying different SOTA methods is an important direction to build a better understanding of self-supervised learning.
Forward citations
Cited by 2 Pith papers
-
Three Necessary Principles for Self-Supervised Visual Representation Learning
Every major self-supervised visual learning method is presented as a special case of a single energy model built from view invariance, spatial prediction, and explicit anti-collapse regularization, which the paper arg...
-
Speaker Verification Under Real Classroom Conditions for English Speech
On a private 18-classroom dataset, WavLM-TDNN with two-stage self-supervised then supervised training achieves the lowest equal error rates among the classroom speaker verification systems tested.
Discussion (0). Continue with ORCID to comment.