Pith. sign in

REVIEW 7 cited by

Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.08981 v2 pith:ETO5LWJN submitted 2021-02-17 cs.CV cs.CL

classification cs.CVcs.CL
keywords pre-trainingconceptualvision-and-languagedatadatasettasksvisualcaptions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The availability of large-scale image captioning and visual question answering datasets has contributed significantly to recent successes in vision-and-language pre-training. However, these datasets are often collected with overrestrictive requirements inherited from their original target tasks (e.g., image caption generation), which limit the resulting dataset scale and diversity. We take a step further in pushing the limits of vision-and-language pre-training data by relaxing the data collection pipeline used in Conceptual Captions 3M (CC3M) [Sharma et al. 2018] and introduce the Conceptual 12M (CC12M), a dataset with 12 million image-text pairs specifically meant to be used for vision-and-language pre-training. We perform an analysis of this dataset and benchmark its effectiveness against CC3M on multiple downstream tasks with an emphasis on long-tail visual recognition. Our results clearly illustrate the benefit of scaling up pre-training data for vision-and-language tasks, as indicated by the new state-of-the-art results on both the nocaps and Conceptual Captions benchmarks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AR-RAG: Autoregressive Retrieval Augmentation for Image Generation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Autoregressive patch-level retrieval augmentation improves text-to-image generation on GenEval, DPG-Bench, and Midjourney-30K, with a training-free decoding variant and a fine-tuned variant.

  2. Native Segmentation Vision Transformers

    cs.CV 2025-05 conditional novelty 7.0 of 10

    A vision transformer backbone that learns to group pixels into semantically coherent segments during downsampling, yielding native segmentation masks without dedicated segmentation heads.

  3. Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models

    cs.RO 2026-02 conditional novelty 6.0 of 10

    Adding a real-world supervised loss to simulation reinforcement learning improves real-robot success and data efficiency for VLA co-training.

  4. SynSeg: Feature Synergy for Multi-Category Contrastive Learning in End-to-End Open-Vocabulary Semantic Segmentation

    cs.CV 2025-08 conditional novelty 6.0 of 10

    SynSeg reports state-of-the-art zero-shot semantic segmentation on four of five benchmarks by combining multi-category contrastive learning with attention-map feature reconstruction.

  5. Info-Coevolution: An Efficient Framework for Data Model Coevolution

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A data-model coevolution framework that fuses model and nearest-neighbor predictions to select labels, reaching ImageNet-1K accuracy with 68% of annotations and 50% under semi-supervised training.

  6. Entity Image and Mixed-Modal Image Retrieval Datasets

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new benchmark, MMIR, built from Wikipedia and WIT by masking entity names and supplying canonical entity images, tests mixed-modal retrieval with single- and multi-entity queries.

  7. TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A training-time pipeline that synthesizes diverse negation captions from batch neighbors improves CLIP's negation accuracy on matching and generation benchmarks, and a new NEG-TTOI benchmark measures negation handling...

Pith tools