Pith. sign in

REVIEW 11 cited by

DataComp: In search of the next generation of multimodal datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.14108 v5 pith:UJAICKRP submitted 2023-04-27 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords datacomptrainingbenchmarkclipbaselinecodecomputedataset
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal datasets are a critical component in recent breakthroughs such as Stable Diffusion and GPT-4, yet their design does not receive the same research attention as model architectures or training algorithms. To address this shortcoming in the ML ecosystem, we introduce DataComp, a testbed for dataset experiments centered around a new candidate pool of 12.8 billion image-text pairs from Common Crawl. Participants in our benchmark design new filtering techniques or curate new data sources and then evaluate their new dataset by running our standardized CLIP training code and testing the resulting model on 38 downstream test sets. Our benchmark consists of multiple compute scales spanning four orders of magnitude, which enables the study of scaling trends and makes the benchmark accessible to researchers with varying resources. Our baseline experiments show that the DataComp workflow leads to better training sets. In particular, our best baseline, DataComp-1B, enables training a CLIP ViT-L/14 from scratch to 79.2% zero-shot accuracy on ImageNet, outperforming OpenAI's CLIP ViT-L/14 by 3.7 percentage points while using the same training procedure and compute. We release DataComp and all accompanying code at www.datacomp.ai.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 74 citations worldwide. Full citation record

  1. RADIO1D: Elastic Representations for Condensed Vision Modeling

    cs.CV 2026-07 accept novelty 7.0 of 10

    RADIO1D produces elastic hierarchical 1D visual tokens via multi-teacher distillation that match or beat fixed 2D encoders in VLMs at lower token counts.

  2. RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement

    cs.SE 2026-07 conditional novelty 6.0 of 10

    A controlled benchmark shows LLM agents can sometimes discover better training-data strategies through feedback, but their improvements are fragile and usually not sustained.

  3. CLIP-like Model as a Foundational Density Ratio Estimator

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CLIP and SigLIP similarity scores can be interpreted as log density ratios, enabling prompt-based importance weighting and KL-divergence-based data curation.

  4. Ambient Diffusion Omni: Training Good Models with Bad Data

    cs.GR 2025-06 conditional novelty 6.0 of 10

    Ambient Diffusion Omni trains diffusion models on mixed-quality data by learning when corrupted images can be treated as clean, improving generation quality and diversity.

  5. CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CuRe scores text-to-image systems by how much their output changes as prompts add cultural details, and reports better agreement with human ratings than existing proxies.

  6. Beginning with You: Perceptual-Initialization Improves Vision-Language Representation and Alignment

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Perceptually initializing a CLIP vision encoder with NIGHTS triplet judgments before YFCC15M contrastive training improves zero-shot accuracy and retrieval over an identical random-start baseline.

  7. Qwen-Audio-VAE Technical Report

    eess.AS 2026-07 conditional novelty 5.0 of 10

    A 12.5 Hz continuous audio VAE reconstructs speech, music, and sound well while encoding 64×30s clips in 541 ms after latency-aware encoder pruning.

  8. MobileCLIP2: Improving Multi-Modal Reinforced Training

    cs.CV 2025-08 conditional novelty 5.0 of 10

    MobileCLIP2 combines DFN-trained teachers, a fine-tuned CoCa captioner, and new 5-stage FastViT variants to set state-of-the-art ImageNet-1k zero-shot accuracy at low latency.

  9. OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning

    cs.CV 2025-09 conditional novelty 4.0 of 10

    OpenVision 2 shows that a caption-only generative objective can match contrastive learning for multimodal vision encoders at lower training cost, scaling to 1B parameters.

  10. An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models

    cs.LG 2025-06 conditional novelty 4.0 of 10

    MultiNet provides an open-source benchmark, data SDK, evaluation harness, and adapted VLA models for assessing generalization across vision, language, and action tasks.

  11. Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences

    cs.CV 2025-06 conditional novelty 4.0 of 10

    SmPO-Diffusion improves diffusion-model preference alignment with reward-model soft labels and ReNoise inversion, reporting higher human-preference scores and up to 26x lower training cost than Diffusion-KTO.

Pith tools