Pith. sign in

REVIEW 14 cited by

Momentum Contrast for Unsupervised Visual Representation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.05722 v3 pith:SE7CS5GL submitted 2019-11-13 cs.CV

classification cs.CV
keywords learningmocounsuperviseddictionaryrepresentationtaskscontrastcontrastive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Momentum Contrast (MoCo) for unsupervised visual representation learning. From a perspective on contrastive learning as dictionary look-up, we build a dynamic dictionary with a queue and a moving-averaged encoder. This enables building a large and consistent dictionary on-the-fly that facilitates contrastive unsupervised learning. MoCo provides competitive results under the common linear protocol on ImageNet classification. More importantly, the representations learned by MoCo transfer well to downstream tasks. MoCo can outperform its supervised pre-training counterpart in 7 detection/segmentation tasks on PASCAL VOC, COCO, and other datasets, sometimes surpassing it by large margins. This suggests that the gap between unsupervised and supervised representation learning has been largely closed in many vision tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 1,012 citations worldwide. Full citation record

  1. Twins: Learn to Predict Unified Representations with Focal Loss

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Channel-wise concatenation of SigLIP2 and Flux VAE features into one token, trained with a focal-style flow-matching loss, yields a unified representation with 1.59 gFID on ImageNet 256 and VAE-level reconstruction.

  2. Bridging the Catalog-to-Real Gap: Scalable Product Recognition via Multi-Stage Contrastive Learning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Multi-stage catalog-to-real contrastive learning (Cat2Real) lifts DINOv3 to 80.73% top-1 accuracy on real-to-catalog product retrieval with strong zero-shot transfer to unseen SKUs and categories.

  3. HASSL: Hierarchy-Aware Self-Supervised Learning Framework for Single Cell Microscopy

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A DINO-style SSL method with a segmentation teacher and stability-weighted HDBSCAN contrastive loss improves hierarchical morphology-aware single-cell embeddings over strong baselines.

  4. Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm

    cs.CV 2025-08 conditional novelty 6.0 of 10

    An audio-conditioned video animation model is pretrained on noisy auto-curated videos and fine-tuned on a few clean examples, achieving top synchronization scores on a new 48-class benchmark with only 1.9% additional ...

  5. SCFlow: Implicitly Learning Style and Content Disentanglement with Flow Models

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    SCFlow learns a reversible style-content merge and then lets the same mapping perform separation without explicit disentanglement training.

  6. Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning

    cs.CV 2025-06 reject novelty 6.0 of 10

    AsymDSD unifies latent masked point modeling and cross-view invariance self-distillation to learn 3D representations, reporting 90.53% on ScanObjectNN and 93.72% with 930k-shape pretraining.

  7. FRAME: Pre-Training Video Feature Representations via Anticipation and Memory

    cs.CV 2025-06 conditional novelty 6.0 of 10

    FRAME distills DINO and CLIP features into a compact video encoder with a memory module and future-frame prediction, outperforming image-based and self-supervised video baselines on dense video tasks.

  8. MABLE: Masked Autoencoding with Bi-Lipschitz Decoding for Embeddings and Graph Metric Learning

    cs.LG 2026-07 conditional novelty 5.5 of 10

    MABLE learns stable node and graph embeddings on heterogeneous geospatial graphs via masked autoencoding, bi-Lipschitz decoding, and fixed cosine alignment/uniformity losses without discriminators or hard negatives.

  9. ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution

    cs.LG 2026-07 conditional novelty 5.0 of 10

    ZUNA1.1, an open-source 380M EEG diffusion autoencoder, reconstructs variable-length, flexibly masked EEG at least as well as its predecessor and far better than spherical spline interpolation.

  10. CIG-MAE: Cross-Modal Information-Guided Masked Autoencoder for Self-Supervised WiFi Sensing

    eess.SP 2025-12 conditional novelty 5.0 of 10

    A dual-stream masked autoencoder with adaptive masking and Barlow Twins alignment learns WiFi CSI representations that beat prior self-supervised baselines and, on SignFi, a fully supervised model.

  11. Barlow-Swin: Toward a novel siamese-based segmentation architecture using Swin-Transformers

    cs.CV 2025-09 reject novelty 4.0 of 10

    Barlow-Swin is a hybrid medical segmenter that pairs a Barlow Twins-pretrained Swin encoder with a U-Net-like decoder, claiming competitive accuracy with fewer parameters.

  12. Foundation Models for Astrophysics

    astro-ph.IM 2026-08 conditional novelty 3.0 of 10

    Astronomical 'foundation models' largely reuse transformers and self-supervised pretraining, but evidence of transfer to new instruments, populations, or tasks remains rare; the paper argues such evidence, not archite...

  13. Acquiring and Adapting Priors for Novel Tasks via Neural Meta-Architectures

    cs.AI 2025-07 conditional novelty 3.0 of 10

    A meta-learning dissertation showing that distributed memory and hypernetworks can adapt to new tasks with few samples, applied to image classification, text-to-3D generation, and molecular binding prediction, with th...

  14. C-LEAD: Contrastive Learning for Enhanced Adversarial Defense

    cs.CV 2025-10 reject novelty 2.0 of 10

    Contrastive learning with adversarial perturbations as positive pairs improves robustness of ResNet models on CIFAR-10, but evidence is weakened by missing baselines and inconsistent reporting.

Pith tools