REVIEW 8 cited by
What If We Recaption Billions of Web Images with LLaMA-3?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Web-crawled image-text pairs are inherently noisy. Prior studies demonstrate that semantically aligning and enriching textual descriptions of these pairs can significantly enhance model training across various vision-language tasks, particularly text-to-image generation. However, large-scale investigations in this area remain predominantly closed-source. Our paper aims to bridge this community effort, leveraging the powerful and \textit{open-sourced} LLaMA-3, a GPT-4 level LLM. Our recaptioning pipeline is simple: first, we fine-tune a LLaMA-3-8B powered LLaVA-1.5 and then employ it to recaption 1.3 billion images from the DataComp-1B dataset. Our empirical results confirm that this enhanced dataset, Recap-DataComp-1B, offers substantial benefits in training advanced vision-language models. For discriminative models like CLIP, we observe enhanced zero-shot performance in cross-modal retrieval tasks. For generative models like text-to-image Diffusion Transformers, the generated images exhibit a significant improvement in alignment with users' text instructions, especially in following complex queries. Our project page is https://www.haqtu.me/Recap-Datacomp-1B/
Forward citations
Cited by 8 Pith papers
-
Seeing More, Saying More: Lightweight Language Experts are Dynamic Video Token Compressors
LangDC compresses video tokens dynamically by converting clips into captions from a small language model, cutting compute by 49% with near-parity accuracy.
-
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models
LVLM-recaptioned image-text data with negative descriptions and short-tag supervision yields a CLIP model that beats larger-data baselines on several benchmarks.
-
LoRA-Loop: Closing the Synthetic Replay Cycle for Continual VLM Learning
Adapting a text-to-image generator with task-specific LoRA adapters and filtering samples by the model's own confidence improves synthetic replay in continual vision-language learning.
-
Advancing Multimodal LLMs by Large-Scale 3D Visual Instruction Dataset Generation
Fine-tuning open vision-language models on 240K synthetic question-answer pairs with exact camera-object labels improves camera-object recognition by 33.4% on average over GPT-4o and Claude-3-Sonnet on the paper's benchmark.
-
Affogato: Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale
A fully automated pipeline using Gemma, Molmo, and SAM generated 750K open-vocabulary 3D affordance annotations on 150K Objaverse objects, and models trained on them transfer to unseen categories.
-
Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models
An image autoencoder compresses pictures into caption-like embeddings via a frozen diffusion decoder, and a fine-tuned LLM reads those embeddings into captions claimed to rival GPT-4o at under $1,000 training cost.
-
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
Structured four-field captions produced small but consistent gains in VQA-based text-image alignment over shuffled versions of the same captions when fine-tuning PixArt-Sigma and Stable Diffusion 2.
-
OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning
OpenVision 2 shows that a caption-only generative objective can match contrastive learning for multimodal vision encoders at lower training cost, scaling to 1B parameters.
Discussion (0). Sign in to comment.