REVIEW 6 cited by
Big Transfer (BiT): General Visual Representation Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Transfer of pre-trained representations improves sample efficiency and simplifies hyperparameter tuning when training deep neural networks for vision. We revisit the paradigm of pre-training on large supervised datasets and fine-tuning the model on a target task. We scale up pre-training, and propose a simple recipe that we call Big Transfer (BiT). By combining a few carefully selected components, and transferring using a simple heuristic, we achieve strong performance on over 20 datasets. BiT performs well across a surprisingly wide range of data regimes -- from 1 example per class to 1M total examples. BiT achieves 87.5% top-1 accuracy on ILSVRC-2012, 99.4% on CIFAR-10, and 76.3% on the 19 task Visual Task Adaptation Benchmark (VTAB). On small datasets, BiT attains 76.8% on ILSVRC-2012 with 10 examples per class, and 97.0% on CIFAR-10 with 10 examples per class. We conduct detailed analysis of the main components that lead to high transfer performance.
Forward citations
Cited by 6 Pith papers
-
Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors
Splitting weight matrices into a fixed-norm direction and learnable per-row/column magnitudes improves LLM training over AdamW/Muon, removes weight decay and warmup, and transfers the optimal LR across width.
-
Open-Set Heterogeneous Domain Adaptation: Theoretical Analysis and Algorithm
RL-OSHeDA, a two-stage representation learning method with pseudo-labeling, outperforms existing domain adaptation baselines on 56 open-set heterogeneous domain adaptation tasks.
-
SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training
SGD is said to minimize free energy, but the temperature is constructed from the data, making the validation circular.
-
Learning Structured Representations with Hyperbolic Embeddings
HypStructure trains image representations whose pairwise distances follow a label tree by adding a hyperbolic CPCC regularizer and a centering loss, reducing hierarchy distortion and modestly improving classification ...
-
Health AI Developer Foundations
Health AI Developer Foundations packages six domain-specific medical embedding models into one platform, claiming large data and compute savings for downstream health ML tasks.
-
Learning from Limited and Imperfect Data
A doctoral thesis compiling nine peer-reviewed papers on long-tailed image generation, long-tailed recognition, semi-supervised learning, and domain adaptation.
Discussion (0). Continue with ORCID to comment.