Pith. sign in

REVIEW 9 cited by

Self-supervised Pretraining of Visual Features in the Wild

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.01988 v2 pith:GE7BCFRL submitted 2021-03-02 cs.CV cs.AI

classification cs.CVcs.AI
keywords self-supervisedlearningrandomdatasetimagenetimagesmethodsmodel
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recently, self-supervised learning methods like MoCo, SimCLR, BYOL and SwAV have reduced the gap with supervised methods. These results have been achieved in a control environment, that is the highly curated ImageNet dataset. However, the premise of self-supervised learning is that it can learn from any random image and from any unbounded dataset. In this work, we explore if self-supervision lives to its expectation by training large models on random, uncurated images with no supervision. Our final SElf-supERvised (SEER) model, a RegNetY with 1.3B parameters trained on 1B random images with 512 GPUs achieves 84.2% top-1 accuracy, surpassing the best self-supervised pretrained model by 1% and confirming that self-supervised learning works in a real world setting. Interestingly, we also observe that self-supervised models are good few-shot learners achieving 77.9% top-1 with access to only 10% of ImageNet. Code: https://github.com/facebookresearch/vissl

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Asymmetric Dual Self-Distillation for 3D Self-Supervised Representation Learning

    cs.CV 2025-06 reject novelty 6.0 of 10

    AsymDSD unifies latent masked point modeling and cross-view invariance self-distillation to learn 3D representations, reporting 90.53% on ScanObjectNN and 93.72% with 930k-shape pretraining.

  2. With Great Backbones Comes Great Adversarial Transferability

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A backbone-only attack that maximizes feature-space distance in a shared pre-trained network transfers to downstream fine-tuned models almost as effectively as white-box attacks.

  3. DINOv2 Meets Text: A Unified Framework for Image- and Pixel-Level Vision-Language Alignment

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Training a text encoder to align with frozen DINOv2 features, plus two small adaptation blocks and balanced data curation, yields strong zero-shot classification and open-vocabulary segmentation at a fraction of CLIP'...

  4. Transferring self-supervised pre-trained models for SHM data anomaly detection with scarce labeled data

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Autoencoder pre-training on unlabeled bridge monitoring data, followed by fine-tuning on a few hundred labels, raises anomaly detection F1 by 3 to 9 points over supervised training in two real bridge datasets.

  5. Pre-training for Action Recognition with Automatically Generated Fractal Datasets

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Automatically generated fractal videos are a viable substitute for Kinetics pre-training in action recognition, matching or exceeding it on two of six benchmarks.

  6. Self-Supervised Radio Pre-training: Toward Foundational Models for Spectrogram Learning

    eess.SP 2024-11 reject novelty 5.0 of 10

    Masked spectrogram pretraining gives radio features that slightly underperform an identical from-scratch baseline on forecasting and clearly underperform on segmentation.

  7. BloomCoreset: Fast Coreset Sampling using Bloom Filters for Fine-Grained Self-Supervised Learning

    cs.CV 2024-12 conditional novelty 4.0 of 10

    BloomCoreset uses Bloom filters on binarized OpenCLIP features to sample image coresets from large open-set pools, cutting sampling time by 98.5% with an average 0.83% accuracy trade-off.

  8. Building 6G Radio Foundation Models with Transformer Architectures

    eess.SP 2024-11 conditional novelty 4.0 of 10

    A masked-autoencoder Vision Transformer pretrained on unlabeled radio spectrograms transfers to human activity sensing and spectrogram segmentation, matching a 4x larger supervised model on the latter.

  9. Instruction-Guided Editing Controls for Images and Multimedia: A Survey in LLM era

    cs.CV 2024-11 unverdicted novelty 3.0 of 10

    A survey that organizes over 100 instruction-guided image and multimedia editing papers into a process-based taxonomy, with an emphasis on LLM and MLLM empowered methods.

Pith tools