Pith. sign in

REVIEW 3 cited by

CLIP Itself is a Strong Fine-tuner: Achieving 85.7% and 88.0% Top-1 Accuracy with ViT-B and ViT-L on ImageNet

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.06138 v1 pith:BJ2NLSZE submitted 2022-12-12 cs.CV cs.LG

classification cs.CVcs.LG
keywords clipfine-tuningperformanceaccuracyhyper-parameteritselftop-1achieve
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent studies have shown that CLIP has achieved remarkable success in performing zero-shot inference while its fine-tuning performance is not satisfactory. In this paper, we identify that fine-tuning performance is significantly impacted by hyper-parameter choices. We examine various key hyper-parameters and empirically evaluate their impact in fine-tuning CLIP for classification tasks through a comprehensive study. We find that the fine-tuning performance of CLIP is substantially underestimated. Equipped with hyper-parameter refinement, we demonstrate CLIP itself is better or at least competitive in fine-tuning compared with large-scale supervised pre-training approaches or latest works that use CLIP as prediction targets in Masked Image Modeling. Specifically, CLIP ViT-Base/16 and CLIP ViT-Large/14 can achieve 85.7%,88.0% finetuning Top-1 accuracy on the ImageNet-1K dataset . These observations challenge the conventional conclusion that CLIP is not suitable for fine-tuning, and motivate us to rethink recently proposed improvements based on CLIP. We will release our code publicly at \url{https://github.com/LightDXY/FT-CLIP}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 16 citations worldwide. Full citation record

  1. FoMo4Wheat: Toward reliable crop vision foundation models with globally curated data

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Training a vision transformer on 2.5 million wheat images outperforms general-domain backbones across ten crop vision tasks.

  2. Mitigating Data Exfiltration Attacks through Layer-Wise Learning Rate Decay Fine-Tuning

    cs.LG 2025-08 conditional novelty 6.0 of 10

    A layer-wise learning rate decay fine-tuning protocol corrupts steganographically embedded training data in exported medical models while preserving classification utility.

  3. Frugal Incremental Generative Modeling using Variational Autoencoders

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A single replay-free conditional VAE with fixed-point-separated Gaussian priors and null-space gradient projection achieves competitive continual classification with drastically reduced memory.

Pith tools