Pith. sign in

REVIEW 6 cited by

Big Transfer (BiT): General Visual Representation Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.11370 v3 pith:UDS5LKDS submitted 2019-12-24 cs.CV cs.LG

classification cs.CVcs.LG
keywords transferclassdatasetsexamplestaskcifar-10componentsilsvrc-2012
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transfer of pre-trained representations improves sample efficiency and simplifies hyperparameter tuning when training deep neural networks for vision. We revisit the paradigm of pre-training on large supervised datasets and fine-tuning the model on a target task. We scale up pre-training, and propose a simple recipe that we call Big Transfer (BiT). By combining a few carefully selected components, and transferring using a simple heuristic, we achieve strong performance on over 20 datasets. BiT performs well across a surprisingly wide range of data regimes -- from 1 example per class to 1M total examples. BiT achieves 87.5% top-1 accuracy on ILSVRC-2012, 99.4% on CIFAR-10, and 76.3% on the 19 task Visual Task Adaptation Benchmark (VTAB). On small datasets, BiT attains 76.8% on ILSVRC-2012 with 10 examples per class, and 97.0% on CIFAR-10 with 10 examples per class. We conduct detailed analysis of the main components that lead to high transfer performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Splitting weight matrices into a fixed-norm direction and learnable per-row/column magnitudes improves LLM training over AdamW/Muon, removes weight decay and warmup, and transfers the optimal LR across width.

  2. Open-Set Heterogeneous Domain Adaptation: Theoretical Analysis and Algorithm

    cs.LG 2024-12 conditional novelty 6.0 of 10

    RL-OSHeDA, a two-stage representation learning method with pseudo-labeling, outperforms existing domain adaptation baselines on 56 open-set heterogeneous domain adaptation tasks.

  3. SGD as Free Energy Minimization: A Thermodynamic View on Neural Network Training

    cs.LG 2025-05 reject novelty 5.0 of 10

    SGD is said to minimize free energy, but the temperature is constructed from the data, making the validation circular.

  4. Learning Structured Representations with Hyperbolic Embeddings

    cs.LG 2024-12 conditional novelty 5.0 of 10

    HypStructure trains image representations whose pairwise distances follow a label tree by adding a hyperbolic CPCC regularizer and a centering loss, reducing hierarchy distortion and modestly improving classification ...

  5. Health AI Developer Foundations

    cs.LG 2024-11 conditional novelty 4.0 of 10

    Health AI Developer Foundations packages six domain-specific medical embedding models into one platform, claiming large data and compute savings for downstream health ML tasks.

  6. Learning from Limited and Imperfect Data

    cs.LG 2025-07 unverdicted novelty 3.0 of 10

    A doctoral thesis compiling nine peer-reviewed papers on long-tailed image generation, long-tailed recognition, semi-supervised learning, and domain adaptation.

Pith tools