Pith. sign in

REVIEW 12 cited by

Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.05918 v2 pith:PMF4WBSR submitted 2021-02-11 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords representationstextvisualdatasetsimagelearningrepresentationvision-language
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-trained representations are becoming crucial for many NLP and perception tasks. While representation learning in NLP has transitioned to training on raw text without human annotations, visual and vision-language representations still rely heavily on curated training datasets that are expensive or require expert knowledge. For vision applications, representations are mostly learned using datasets with explicit class labels such as ImageNet or OpenImages. For vision-language, popular datasets like Conceptual Captions, MSCOCO, or CLIP all involve a non-trivial data collection (and cleaning) process. This costly curation process limits the size of datasets and hence hinders the scaling of trained models. In this paper, we leverage a noisy dataset of over one billion image alt-text pairs, obtained without expensive filtering or post-processing steps in the Conceptual Captions dataset. A simple dual-encoder architecture learns to align visual and language representations of the image and text pairs using a contrastive loss. We show that the scale of our corpus can make up for its noise and leads to state-of-the-art representations even with such a simple learning scheme. Our visual representation achieves strong performance when transferred to classification tasks such as ImageNet and VTAB. The aligned visual and language representations enables zero-shot image classification and also set new state-of-the-art results on Flickr30K and MSCOCO image-text retrieval benchmarks, even when compared with more sophisticated cross-attention models. The representations also enable cross-modality search with complex text and text + image queries.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 1,196 citations worldwide. Full citation record

  1. WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition

    cs.CV 2026-03 unverdicted novelty 7.0 of 10

    WikiCLIP reaches 28.5% OVEN-unseen accuracy (vs 24.5% AutoVER) at 14.5 ms latency by vision-guided LLM embeddings plus hard-negative text swaps.

  2. The Vision Wormhole: Latent-Space Communication in Heterogeneous Multi-Agent Systems

    cs.CL 2026-02 conditional novelty 6.0 of 10

    Reasoning messages between heterogeneous VLMs can be routed through the image-token span: a distilled universal codec plus affine alignment transmits latent traces across model families, cutting wall-clock time in sma...

  3. Rate-Distortion Limits for Multimodal Retrieval: Theory, Optimal Codes, and Finite-Sample Guarantees

    cs.IT 2025-09 reject novelty 6.0 of 10

    A claimed rate-distortion limit for multimodal retrieval with a modality-skew penalty and an adaptive-temperature quantizer, but the supporting derivation is invalid.

  4. VLMs-in-the-Wild: Bridging the Gap Between Academic Benchmarks and Enterprise Reality

    cs.CV 2025-09 conditional novelty 6.0 of 10

    VLM-in-the-Wild provides an enterprise-focused benchmark and the BlockWeaver OCR matching algorithm, reporting that a small fine-tuned model can rival a 32B model on some tasks.

  5. Visual Pre-Training on Unlabeled Images using Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Casting image-crop consistency as temporal-difference value learning improves visual representations on unlabeled web, scene, and video data.

  6. WisWheat: A Three-Tiered Vision-Language Dataset for Wheat Management

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new wheat-specific dataset with pretraining, quantitative, and instruction-tuning layers improves VLM performance on wheat stress diagnosis and growth-stage management tasks.

  7. A Real-to-Sim-to-Real Approach to Robotic Manipulation with VLM-Generated Iterative Keypoint Rewards

    cs.RO 2025-02 conditional novelty 6.0 of 10

    IKER uses VLM-generated keypoint rewards to train manipulation policies in simulation that transfer to a real robot, enabling multi-step tasks and replanning.

  8. Theia: Large-Scale Multimodal Captioning and Automated Validation of the Incidents1M Dataset for Data-Free Distillation

    cs.CV 2026-07 conditional novelty 5.5 of 10

    Incidents1M is extended with 200k VLM captions and validated by an image-blind LLM-as-a-Judge that reports high inter-model agreement and conservative high-precision labeling.

  9. Qwen-Audio-VAE Technical Report

    eess.AS 2026-07 conditional novelty 5.0 of 10

    A 12.5 Hz continuous audio VAE reconstructs speech, music, and sound well while encoding 64×30s clips in 541 ms after latency-aware encoder pruning.

  10. Robust and Label-Efficient Deep Waste Detection

    cs.CV 2025-08 conditional novelty 5.0 of 10

    An ensemble-based soft pseudo-labeling pipeline improves waste detection on the ZeroWaste dataset, beating fully supervised training with the same labeled images.

  11. CF-VLM:CounterFactual Vision-Language Fine-tuning

    cs.LG 2025-06 conditional novelty 5.0 of 10

    CF-VLM fine-tunes VLMs on counterfactual image-text pairs with three objectives, reporting gains on compositional reasoning benchmarks and modest hallucination reductions.

  12. A Survey on Training-free Open-Vocabulary Semantic Segmentation

    cs.CV 2025-05 conditional novelty 2.0 of 10

    A structured review of over 30 training-free open-vocabulary semantic segmentation methods, organized by whether they rely on CLIP alone, auxiliary visual foundation models, or generative models.

Pith tools