REVIEW 6 cited by
Big Self-Supervised Models are Strong Semi-Supervised Learners
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
abstract
One paradigm for learning from few labeled examples while making best use of a large amount of unlabeled data is unsupervised pretraining followed by supervised fine-tuning. Although this paradigm uses unlabeled data in a task-agnostic way, in contrast to common approaches to semi-supervised learning for computer vision, we show that it is surprisingly effective for semi-supervised learning on ImageNet. A key ingredient of our approach is the use of big (deep and wide) networks during pretraining and fine-tuning. We find that, the fewer the labels, the more this approach (task-agnostic use of unlabeled data) benefits from a bigger network. After fine-tuning, the big network can be further improved and distilled into a much smaller one with little loss in classification accuracy by using the unlabeled examples for a second time, but in a task-specific way. The proposed semi-supervised learning algorithm can be summarized in three steps: unsupervised pretraining of a big ResNet model using SimCLRv2, supervised fine-tuning on a few labeled examples, and distillation with unlabeled examples for refining and transferring the task-specific knowledge. This procedure achieves 73.9% ImageNet top-1 accuracy with just 1% of the labels ($\le$13 labeled images per class) using ResNet-50, a $10\times$ improvement in label efficiency over the previous state-of-the-art. With 10% of labels, ResNet-50 trained with our method achieves 77.5% top-1 accuracy, outperforming standard supervised training with all of the labels.
Forward citations
Cited by 6 Pith papers
-
Visual Pre-Training on Unlabeled Images using Reinforcement Learning
Casting image-crop consistency as temporal-difference value learning improves visual representations on unlabeled web, scene, and video data.
-
Examining the Efficacy of Graph Neural Network Message-Passing in Regression Contexts
Across four NAS/DNN-predictor regression benchmarks, GEN (deep graph convolution) achieves the best average rank over 11 GNN message-passing layers, though attention GATv2 wins on the largest graphs.
-
Long-Tailed Object Detection Pre-training: Dynamic Rebalancing Contrastive Learning with Dual Reconstruction
A detection pre-training framework that dynamically rebalances rare classes and adds dual reconstruction improves tail-class AP on COCO and LVIS by about 0.5 to 1.5 points in most configurations.
-
Self-Supervised Learning for Pre-training Capsule Networks: Overcoming Medical Imaging Dataset Challenges
Contrastive learning with in-painting as a self-supervised pre-training task slightly improves capsule network accuracy on PICCOLO polyp classification (0.40 vs 0.38) but gives worse AUROC than ImageNet pre-training.
-
Contrastive Masked Autoencoders for Character-Level Open-Set Writer Identification
A contrastive masked autoencoder achieves 89.7% precision on open-set character-level writer identification on CASIA-OLHWDB, and 81.6% rank-1 on IAM-OnDB.
-
Object Tracking in a $360^o$ View: A Novel Perspective on Bridging the Gap to Biomedical Advancements
A review of object tracking algorithms for biomedical video concludes that deep learning is the most capable family, but it includes a placeholder citation for a model described as real.
Discussion (0). Continue with ORCID to comment.