Pith. sign in

REVIEW 3 cited by

A Hard-to-Beat Baseline for Training-free CLIP-based Adaptation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.04087 v1 pith:77XMKQSJ submitted 2024-02-06 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords cliplearningmethodsclassclassificationclassifiercovariancedownstream
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Contrastive Language-Image Pretraining (CLIP) has gained popularity for its remarkable zero-shot capacity. Recent research has focused on developing efficient fine-tuning methods, such as prompt learning and adapter, to enhance CLIP's performance in downstream tasks. However, these methods still require additional training time and computational resources, which is undesirable for devices with limited resources. In this paper, we revisit a classical algorithm, Gaussian Discriminant Analysis (GDA), and apply it to the downstream classification of CLIP. Typically, GDA assumes that features of each class follow Gaussian distributions with identical covariance. By leveraging Bayes' formula, the classifier can be expressed in terms of the class means and covariance, which can be estimated from the data without the need for training. To integrate knowledge from both visual and textual modalities, we ensemble it with the original zero-shot classifier within CLIP. Extensive results on 17 datasets validate that our method surpasses or achieves comparable results with state-of-the-art methods on few-shot classification, imbalanced learning, and out-of-distribution generalization. In addition, we extend our method to base-to-new generalization and unsupervised learning, once again demonstrating its superiority over competing approaches. Our code is publicly available at \url{https://github.com/mrflogs/ICLR24}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Von Mises-Fisher Mixture Model with Dynamic Shrinkage for Realistic Test-Time Transduction

    cs.CV 2026-07 conditional novelty 6.0 of 10

    KL-anchored penalized likelihood with class- and instance-dependent shrinkage, implemented with von Mises-Fisher mixtures, improves CLIP test-time transduction under class imbalance.

  2. Simulate, Refocus and Ensemble: An Attention-Refocusing Scheme for Domain Generalization

    cs.CV 2025-07 conditional novelty 6.0 of 10

    SRE improves CLIP's domain generalization by training an attention-refocuser on simulated target domains and ensembling the most attention-consistent checkpoints.

  3. Generalizing vision-language models to novel domains: A comprehensive survey

    cs.CV 2025-06 conditional novelty 3.0 of 10

    A survey of VLM generalization literature organized by transferred module, with benchmark tables and a review of multimodal LLMs.

Pith tools