REVIEW 9 cited by
Unsupervised Learning of Visual Features by Contrasting Cluster Assignments
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Unsupervised image representations have significantly reduced the gap with supervised pretraining, notably with the recent achievements of contrastive learning methods. These contrastive methods typically work online and rely on a large number of explicit pairwise feature comparisons, which is computationally challenging. In this paper, we propose an online algorithm, SwAV, that takes advantage of contrastive methods without requiring to compute pairwise comparisons. Specifically, our method simultaneously clusters the data while enforcing consistency between cluster assignments produced for different augmentations (or views) of the same image, instead of comparing features directly as in contrastive learning. Simply put, we use a swapped prediction mechanism where we predict the cluster assignment of a view from the representation of another view. Our method can be trained with large and small batches and can scale to unlimited amounts of data. Compared to previous contrastive methods, our method is more memory efficient since it does not require a large memory bank or a special momentum network. In addition, we also propose a new data augmentation strategy, multi-crop, that uses a mix of views with different resolutions in place of two full-resolution views, without increasing the memory or compute requirements much. We validate our findings by achieving 75.3% top-1 accuracy on ImageNet with ResNet-50, as well as surpassing supervised pretraining on all the considered transfer tasks.
Forward citations
Cited by 9 Pith papers
-
NSD-Imagery: A benchmark dataset for extending fMRI vision decoding methods to mental imagery
NSD-Imagery is a released benchmark of fMRI responses to imagined pictures from Natural Scenes Dataset participants, and benchmarks of five decoders show mental imagery performance is largely decoupled from seen-image...
-
A Cross Branch Fusion-Based Contrastive Learning Framework for Point Cloud Self-supervised Learning
PoCCA improves point cloud self-supervised learning by fusing online and target branch features via cross-attention before the contrastive loss, achieving state-of-the-art among methods without extra training data.
-
Real-time Reconstruction of Human Visual Perception from fMRI
First demonstration that single-trial visual images can be decoded from fMRI in near-real-time (about 10-15 seconds) with roughly one hour of fine-tuning data.
-
Visual Pre-Training on Unlabeled Images using Reinforcement Learning
Casting image-crop consistency as temporal-difference value learning improves visual representations on unlabeled web, scene, and video data.
-
VET-DINO: Learning Anatomical Understanding Through Multi-View Distillation in Veterinary Imaging
Using real multi-view radiographs from the same study as self-supervised training pairs yields better anatomical representations and downstream veterinary task performance than synthetic single-image augmentations.
-
A Generalized Learning Framework for Self-Supervised Contrastive Learning
A single framework unifies BYOL, Barlow Twins, and SwAV, plus a plug-in calibration method, ADC, that improves learned representations by preserving input-space distances.
-
On the Out-of-Distribution Generalization of Self-Supervised Learning
Self-supervised learning can be made more robust to distribution shift by sampling mini-batches so that spurious background variables are independent of the anchor label, using a VAE and balancing-score matching.
-
Barlow-Swin: Toward a novel siamese-based segmentation architecture using Swin-Transformers
Barlow-Swin is a hybrid medical segmenter that pairs a Barlow Twins-pretrained Swin encoder with a U-Net-like decoder, claiming competitive accuracy with fewer parameters.
-
Quantum Feature Optimization for Enhanced Clustering of Blockchain Transaction Data
Quantum feature maps are reported to improve blockchain transaction clustering, but the comparison omits classical random features and the results are selected on the test set.
Discussion (0). Continue with ORCID to comment.