REVIEW 7 cited by
Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The availability of large-scale image captioning and visual question answering datasets has contributed significantly to recent successes in vision-and-language pre-training. However, these datasets are often collected with overrestrictive requirements inherited from their original target tasks (e.g., image caption generation), which limit the resulting dataset scale and diversity. We take a step further in pushing the limits of vision-and-language pre-training data by relaxing the data collection pipeline used in Conceptual Captions 3M (CC3M) [Sharma et al. 2018] and introduce the Conceptual 12M (CC12M), a dataset with 12 million image-text pairs specifically meant to be used for vision-and-language pre-training. We perform an analysis of this dataset and benchmark its effectiveness against CC3M on multiple downstream tasks with an emphasis on long-tail visual recognition. Our results clearly illustrate the benefit of scaling up pre-training data for vision-and-language tasks, as indicated by the new state-of-the-art results on both the nocaps and Conceptual Captions benchmarks.
Forward citations
Cited by 7 Pith papers
-
AR-RAG: Autoregressive Retrieval Augmentation for Image Generation
Autoregressive patch-level retrieval augmentation improves text-to-image generation on GenEval, DPG-Bench, and Midjourney-30K, with a training-free decoding variant and a fine-tuned variant.
-
Native Segmentation Vision Transformers
A vision transformer backbone that learns to group pixels into semantically coherent segments during downsampling, yielding native segmentation masks without dedicated segmentation heads.
-
Beyond Imitation: Reinforcement Learning-Based Sim-Real Co-Training for VLA Models
Adding a real-world supervised loss to simulation reinforcement learning improves real-robot success and data efficiency for VLA co-training.
-
SynSeg: Feature Synergy for Multi-Category Contrastive Learning in End-to-End Open-Vocabulary Semantic Segmentation
SynSeg reports state-of-the-art zero-shot semantic segmentation on four of five benchmarks by combining multi-category contrastive learning with attention-map feature reconstruction.
-
Info-Coevolution: An Efficient Framework for Data Model Coevolution
A data-model coevolution framework that fuses model and nearest-neighbor predictions to select labels, reaching ImageNet-1K accuracy with 68% of annotations and 50% under semi-supervised training.
-
Entity Image and Mixed-Modal Image Retrieval Datasets
A new benchmark, MMIR, built from Wikipedia and WIT by masking entity names and supplying canonical entity images, tests mixed-modal retrieval with single- and multi-entity queries.
-
TNG-CLIP:Training-Time Negation Data Generation for Negation Awareness of CLIP
A training-time pipeline that synthesizes diverse negation captions from batch neighbors improves CLIP's negation accuracy on matching and generation benchmarks, and a new NEG-TTOI benchmark measures negation handling...
Discussion (0). Continue with ORCID to comment.