REVIEW 14 cited by
Momentum Contrast for Unsupervised Visual Representation Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present Momentum Contrast (MoCo) for unsupervised visual representation learning. From a perspective on contrastive learning as dictionary look-up, we build a dynamic dictionary with a queue and a moving-averaged encoder. This enables building a large and consistent dictionary on-the-fly that facilitates contrastive unsupervised learning. MoCo provides competitive results under the common linear protocol on ImageNet classification. More importantly, the representations learned by MoCo transfer well to downstream tasks. MoCo can outperform its supervised pre-training counterpart in 7 detection/segmentation tasks on PASCAL VOC, COCO, and other datasets, sometimes surpassing it by large margins. This suggests that the gap between unsupervised and supervised representation learning has been largely closed in many vision tasks.
Forward citations
Cited by 14 Pith papers
-
Twins: Learn to Predict Unified Representations with Focal Loss
Channel-wise concatenation of SigLIP2 and Flux VAE features into one token, trained with a focal-style flow-matching loss, yields a unified representation with 1.59 gFID on ImageNet 256 and VAE-level reconstruction.
-
Bridging the Catalog-to-Real Gap: Scalable Product Recognition via Multi-Stage Contrastive Learning
Multi-stage catalog-to-real contrastive learning (Cat2Real) lifts DINOv3 to 80.73% top-1 accuracy on real-to-catalog product retrieval with strong zero-shot transfer to unseen SKUs and categories.
-
HASSL: Hierarchy-Aware Self-Supervised Learning Framework for Single Cell Microscopy
A DINO-style SSL method with a segmentation teacher and stability-weighted HDBSCAN contrastive loss improves hierarchical morphology-aware single-cell embeddings over strong baselines.
-
Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm
An audio-conditioned video animation model is pretrained on noisy auto-curated videos and fine-tuned on a few clean examples, achieving top synchronization scores on a new 48-class benchmark with only 1.9% additional ...
-
SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models
SCFlow learns a reversible style-content merge and then lets the same mapping perform separation without explicit disentanglement training.
-
Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning
AsymDSD unifies latent masked point modeling and cross-view invariance self-distillation to learn 3D representations, reporting 90.53% on ScanObjectNN and 93.72% with 930k-shape pretraining.
-
FRAME: Pre-Training Video Feature Representations via Anticipation and Memory
FRAME distills DINO and CLIP features into a compact video encoder with a memory module and future-frame prediction, outperforming image-based and self-supervised video baselines on dense video tasks.
-
MABLE: Masked Autoencoding with Bi-Lipschitz Decoding for Embeddings and Graph Metric Learning
MABLE learns stable node and graph embeddings on heterogeneous geospatial graphs via masked autoencoding, bi-Lipschitz decoding, and fixed cosine alignment/uniformity losses without discriminators or hard negatives.
-
ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution
ZUNA1.1, an open-source 380M EEG diffusion autoencoder, reconstructs variable-length, flexibly masked EEG at least as well as its predecessor and far better than spherical spline interpolation.
-
CIG-MAE: Cross-Modal Information-Guided Masked Autoencoder for Self-Supervised WiFi Sensing
A dual-stream masked autoencoder with adaptive masking and Barlow Twins alignment learns WiFi CSI representations that beat prior self-supervised baselines and, on SignFi, a fully supervised model.
-
Barlow-Swin: Toward a novel siamese-based segmentation architecture using Swin-Transformers
Barlow-Swin is a hybrid medical segmenter that pairs a Barlow Twins-pretrained Swin encoder with a U-Net-like decoder, claiming competitive accuracy with fewer parameters.
-
Foundation Models for Astrophysics
Astronomical 'foundation models' largely reuse transformers and self-supervised pretraining, but evidence of transfer to new instruments, populations, or tasks remains rare; the paper argues such evidence, not archite...
-
Acquiring and Adapting Priors for Novel Tasks via Neural Meta-Architectures
A meta-learning dissertation showing that distributed memory and hypernetworks can adapt to new tasks with few samples, applied to image classification, text-to-3D generation, and molecular binding prediction, with th...
-
C-LEAD: Contrastive Learning for Enhanced Adversarial Defense
Contrastive learning with adversarial perturbations as positive pairs improves robustness of ResNet models on CIFAR-10, but evidence is weakened by missing baselines and inconsistent reporting.
Discussion (0). Sign in to comment.