Pith. sign in

REVIEW 8 cited by

What If We Recaption Billions of Web Images with LLaMA-3?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.08478 v2 pith:4FS3OBEF submitted 2024-06-12 cs.CV cs.CL

classification cs.CVcs.CL
keywords imagesmodelsdatasetenhancedlikellama-3pairsrecap-datacomp-1b
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Web-crawled image-text pairs are inherently noisy. Prior studies demonstrate that semantically aligning and enriching textual descriptions of these pairs can significantly enhance model training across various vision-language tasks, particularly text-to-image generation. However, large-scale investigations in this area remain predominantly closed-source. Our paper aims to bridge this community effort, leveraging the powerful and \textit{open-sourced} LLaMA-3, a GPT-4 level LLM. Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered LLaVA-1.5 and then employ it to recaption 1.3 billion images from the DataComp-1B dataset. Our empirical results confirm that this enhanced dataset, Recap-DataComp-1B, offers substantial benefits in training advanced vision-language models. For discriminative models like CLIP, we observe enhanced zero-shot performance in cross-modal retrieval tasks. For generative models like text-to-image Diffusion Transformers, the generated images exhibit a significant improvement in alignment with users' text instructions, especially in following complex queries. Our project page is https://www.haqtu.me/Recap-Datacomp-1B/

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors

    cs.CV 2025-08 conditional novelty 6.0 of 10

    LangDC compresses video tokens dynamically by converting clips into captions from a small language model, cutting compute by 49% with near-parity accuracy.

  2. HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LVLM-recaptioned image-text data with negative descriptions and short-tag supervision yields a CLIP model that beats larger-data baselines on several benchmarks.

  3. LoRA-Loop: Closing the Synthetic Replay Cycle for Continual VLM Learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Adapting a text-to-image generator with task-specific LoRA adapters and filtering samples by the model's own confidence improves synthetic replay in continual vision-language learning.

  4. Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation

    cs.GR 2025-07 conditional novelty 6.0 of 10

    Fine-tuning open vision-language models on 240K synthetic question-answer pairs with exact camera-object labels improves camera-object recognition by 33.4% on average over GPT-4o and Claude-3-Sonnet on the paper's benchmark.

  5. Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A fully automated pipeline using Gemma, Molmo, and SAM generated 750K open-vocabulary 3D affordance annotations on 150K Objaverse objects, and models trained on them transfer to unseen categories.

  6. Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models

    cs.CV 2025-07 conditional novelty 5.0 of 10

    An image autoencoder compresses pictures into caption-like embeddings via a frozen diffusion decoder, and a fine-tuned LLM reads those embeddings into captions claimed to rival GPT-4o at under $1,000 training cost.

  7. Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Structured four-field captions produced small but consistent gains in VQA-based text-image alignment over shuffled versions of the same captions when fine-tuning PixArt-Sigma and Stable Diffusion 2.

  8. OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning

    cs.CV 2025-09 conditional novelty 4.0 of 10

    OpenVision 2 shows that a caption-only generative objective can match contrastive learning for multimodal vision encoders at lower training cost, scaling to 1B parameters.

Pith tools