Pith. sign in

REVIEW 3 cited by

No "Zero-Shot" Without Exponential Data: Pretraining Concept Frequency Determines Multimodal Model Performance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.04125 v3 pith:6TKQTDA2 submitted 2024-04-04 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords modelszero-shotdatamultimodalpretrainingdatasetsdownstreamperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Web-crawled pretraining datasets underlie the impressive "zero-shot" evaluation performance of multimodal models, such as CLIP for classification/retrieval and Stable-Diffusion for image generation. However, it is unclear how meaningful the notion of "zero-shot" generalization is for such multimodal models, as it is not known to what extent their pretraining datasets encompass the downstream concepts targeted for during "zero-shot" evaluation. In this work, we ask: How is the performance of multimodal models on downstream concepts influenced by the frequency of these concepts in their pretraining datasets? We comprehensively investigate this question across 34 models and five standard pretraining datasets (CC-3M, CC-12M, YFCC-15M, LAION-400M, LAION-Aesthetics), generating over 300GB of data artifacts. We consistently find that, far from exhibiting "zero-shot" generalization, multimodal models require exponentially more data to achieve linear improvements in downstream "zero-shot" performance, following a sample inefficient log-linear scaling trend. This trend persists even when controlling for sample-level similarity between pretraining and downstream datasets, and testing on purely synthetic data distributions. Furthermore, upon benchmarking models on long-tailed data sampled based on our analysis, we demonstrate that multimodal models across the board perform poorly. We contribute this long-tail test set as the "Let it Wag!" benchmark to further research in this direction. Taken together, our study reveals an exponential need for training data which implies that the key to "zero-shot" generalization capabilities under large-scale training paradigms remains to be found.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Modality Fusion: Deep Ensembles for Multimodal Classification

    cs.LG 2026-07 accept novelty 6.5 of 10

    Heterogeneous deep ensembles of unimodal models outperform imbalance-aware late-fusion and intermediate-fusion multimodal classifiers at matched parameter count, with a loss-ratio rule for allocating models per modality.

  2. Prompt Tuning Vision Language Models with Margin Regularizer for Few-Shot Learning under Distribution Shifts

    cs.CV 2025-05 conditional novelty 5.0 of 10

    PromptMargin adapts CLIP to few-shot classification under distribution shift using selective augmentations and a multimodal margin regularizer, beating MaPLe on most of fifteen datasets.

  3. Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget

    cs.CV 2026-07 conditional novelty 4.0 of 10

    An open-source image-generation family shows that agentic prompt rewriting and a stronger text encoder can lift quality to near closed-source levels with only 208.62M images and about $400K of training compute.

Pith tools