Pith. sign in

REVIEW 6 cited by

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1908.02265 v1 pith:J4HQU52O submitted 2019-08-06 cs.CV cs.CL

classification cs.CVcs.CL
keywords tasksvisualmodelvision-and-languagearchitecturebertimagelanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present ViLBERT (short for Vision-and-Language BERT), a model for learning task-agnostic joint representations of image content and natural language. We extend the popular BERT architecture to a multi-modal two-stream model, pro-cessing both visual and textual inputs in separate streams that interact through co-attentional transformer layers. We pretrain our model through two proxy tasks on the large, automatically collected Conceptual Captions dataset and then transfer it to multiple established vision-and-language tasks -- visual question answering, visual commonsense reasoning, referring expressions, and caption-based image retrieval -- by making only minor additions to the base architecture. We observe significant improvements across tasks compared to existing task-specific models -- achieving state-of-the-art on all four tasks. Our work represents a shift away from learning groundings between vision and language only as part of task training and towards treating visual grounding as a pretrainable and transferable capability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 1,675 citations worldwide. Full citation record

  1. Representations in vision and language converge in a shared, multidimensional space of perceived similarities

    q-bio.NC 2025-07 conditional novelty 6.0 of 10

    Similarity judgments of natural scene images and their sentence captions are aligned with each other, with visual brain responses, and with LLM-trained visual models.

  2. LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    LaVi encodes visual context into LayerNorm affine parameters, bypassing visual token concatenation, and reports LLaVA-comparable accuracy at a 94% FLOP reduction.

  3. From Image Captioning to Visual Storytelling

    cs.CL 2025-07 unverdicted novelty 4.0 of 10

    Visual storytelling improves by treating it as image captioning followed by language-to-language story generation, with a new 'ideality' metric to gauge distance from an oracle.

  4. On the Resilience of Underwater Semantic Wireless Communications

    cs.NI 2025-06 conditional novelty 4.0 of 10

    In a simulated underwater acoustic link, the SAGE semantic image system keeps semantic similarity around 50% up to 15-20% character error, indicating resilience to text corruption.

  5. Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval

    cs.CV 2025-05 conditional novelty 4.0 of 10

    RDB improves remote sensing image-text retrieval mean recall by 1.15 to 2 percent over fully fine-tuned GeoRSCLIP using an asymmetric adapter and a dual-task consistency loss.

  6. Vision-Language Models for Edge Networks: A Comprehensive Survey

    cs.CV 2025-02 reject novelty 2.0 of 10

    A survey of lightweight vision-language models for edge deployment, marred by citation errors, self-citation, and a lack of selection methodology.

Pith tools