Pith. sign in

REVIEW 4 cited by

OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.16527 v2 pith:R5QWLZJH submitted 2023-06-21 cs.IR cs.CV

classification cs.IRcs.CV
keywords datasetmodelsdocumentsimage-textmultimodalobelicsbeenbenchmarks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large multimodal models trained on natural documents, which interleave images and text, outperform models trained on image-text pairs on various multimodal benchmarks. However, the datasets used to train these models have not been released, and the collection process has not been fully specified. We introduce the OBELICS dataset, an open web-scale filtered dataset of interleaved image-text documents comprising 141 million web pages extracted from Common Crawl, 353 million associated images, and 115 billion text tokens. We describe the dataset creation process, present comprehensive filtering rules, and provide an analysis of the dataset's content. To show the viability of OBELICS, we train vision and language models of 9 and 80 billion parameters named IDEFICS, and obtain competitive performance on different multimodal benchmarks. We release our dataset, models and code.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A benchmark that scores vision-language models by reconstructing the 3D scene behind an image as executable Blender code finds the models fail mainly on spatial precision, not tool usage.

  2. AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A finetuned vision-language model jointly predicts nine aspect scores and written comments for AI-generated videos, with a new benchmark and claims of state-of-the-art alignment with human judgment.

  3. BcQLM: Efficient Vision-Language Understanding with Distilled Q-Gated Cross-Modal Fusion

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A roughly 1.2B-parameter VQA model with a distilled 31M CLIP encoder and Q-gated cross-attention reports accuracies comparable to 7B-13B baselines on GQA, VQAv2, and VizWiz.

  4. An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models

    cs.LG 2025-06 conditional novelty 4.0 of 10

    MultiNet provides an open-source benchmark, data SDK, evaluation harness, and adapted VLA models for assessing generalization across vision, language, and action tasks.

Pith tools