REVIEW 4 cited by
OBELICS: An Open Web-Scale Filtered Dataset of Interleaved Image-Text Documents
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large multimodal models trained on natural documents, which interleave images and text, outperform models trained on image-text pairs on various multimodal benchmarks. However, the datasets used to train these models have not been released, and the collection process has not been fully specified. We introduce the OBELICS dataset, an open web-scale filtered dataset of interleaved image-text documents comprising 141 million web pages extracted from Common Crawl, 353 million associated images, and 115 billion text tokens. We describe the dataset creation process, present comprehensive filtering rules, and provide an analysis of the dataset's content. To show the viability of OBELICS, we train vision and language models of 9 and 80 billion parameters named IDEFICS, and obtain competitive performance on different multimodal benchmarks. We release our dataset, models and code.
Forward citations
Cited by 4 Pith papers
-
IR3D-Bench: Evaluating Vision-Language Model Scene Understanding as Agentic Inverse Rendering
A benchmark that scores vision-language models by reconstructing the 3D scene behind an image as executable Blender code finds the models fail mainly on spatial precision, not tool usage.
-
AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation
A finetuned vision-language model jointly predicts nine aspect scores and written comments for AI-generated videos, with a new benchmark and claims of state-of-the-art alignment with human judgment.
-
BcQLM: Efficient Vision-Language Understanding with Distilled Q-Gated Cross-Modal Fusion
A roughly 1.2B-parameter VQA model with a distilled 31M CLIP encoder and Q-gated cross-attention reports accuracies comparable to 7B-13B baselines on GQA, VQAv2, and VizWiz.
-
An Open-Source Software Toolkit & Benchmark Suite for the Evaluation and Adaptation of Multimodal Action Models
MultiNet provides an open-source benchmark, data SDK, evaluation harness, and adapted VLA models for assessing generalization across vision, language, and action tasks.
Discussion (0). Sign in to comment.