Pith. sign in

REVIEW 11 cited by

VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1908.08530 v4 pith:ZYEKBKGB submitted 2019-08-22 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords vl-bertvisual-linguisticinputgenerictasksvisualbetterdownstream
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce a new pre-trainable generic representation for visual-linguistic tasks, called Visual-Linguistic BERT (VL-BERT for short). VL-BERT adopts the simple yet powerful Transformer model as the backbone, and extends it to take both visual and linguistic embedded features as input. In it, each element of the input is either of a word from the input sentence, or a region-of-interest (RoI) from the input image. It is designed to fit for most of the visual-linguistic downstream tasks. To better exploit the generic representation, we pre-train VL-BERT on the massive-scale Conceptual Captions dataset, together with text-only corpus. Extensive empirical analysis demonstrates that the pre-training procedure can better align the visual-linguistic clues and benefit the downstream tasks, such as visual commonsense reasoning, visual question answering and referring expression comprehension. It is worth noting that VL-BERT achieved the first place of single model on the leaderboard of the VCR benchmark. Code is released at \url{https://github.com/jackroos/VL-BERT}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Relational Tabular Data without Shared Features

    cs.LG 2025-02 reject novelty 7.0 of 10

    Leal learns cross-table alignment without shared features by treating lower training loss as evidence of correct row matching, with a cluster sampler to limit the candidate set.

  2. PVCap: Towards Accurate 3D Dense Captioning via PseudoCap and VoxelCapNet

    cs.CV 2026-07 conditional novelty 6.0 of 10

    PVCap combines instance-mixing data augmentation with pseudo-labels and a voxel-based captioning network to achieve new state-of-the-art on 3D dense captioning benchmarks ScanRefer and Nr3D.

  3. L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A distilled 99M-parameter CLIP gives L-CLIPScore, a lightweight caption metric that matches CLIPScore on human correlation and best improves captioning models when mixed with CIDEr.

  4. Biases Propagate in Encoder-based Vision-Language Models: A Systematic Analysis From Intrinsic Measures to Zero-shot Retrieval Outcomes

    cs.CL 2025-06 reject novelty 6.0 of 10

    Intrinsic SC-EAT bias scores in CLIP and BLIP-2 embeddings correlate strongly (average Spearman rho = 0.83) with biased zero-shot retrieval outcomes, but the shared-stimulus design may inflate this correlation.

  5. Generative AI for Testing of Autonomous Driving Systems: A Survey

    cs.SE 2025-08 conditional novelty 5.0 of 10

    A systematic survey that organizes 91 studies of generative AI for autonomous driving testing into six scenario-based tasks and catalogs 27 limitations.

  6. Computed Tomography Visual Question Answering with Cross-modal Feature Graphing

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A cross-modal graph connecting CT slices and question tokens, aggregated by an attentive GCN, improves LLM-based CT visual question answering on M3D-VQA.

  7. Multimodal Prompt Alignment for Facial Expression Recognition

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A frozen-CLIP facial expression recognition framework with LLM-guided prompts, prototype regularization, and sparse global-local alignment claims new state-of-the-art accuracy across RAF-DB, FERPlus, and AffectNet.

  8. MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Grouping instruction-tuning datasets by redundancy, uniqueness, or synergy of text-image interaction improves vision-language model accuracy over single-task and unselective multi-task tuning.

  9. Kronecker Mask and Interpretive Prompts are Language-Action Video Learners

    cs.CV 2025-02 conditional novelty 5.0 of 10

    CLAVER adds a cross-frame temporal attention mask (Kronecker mask) and LLM-generated interpretive action prompts to CLIP, improving video action recognition.

  10. Transformer-based Spatial Grounding: A Comprehensive Survey

    cs.CV 2025-07 reject novelty 4.0 of 10

    A systematic literature review of 45 papers on transformer-based spatial grounding catalogs architectures, datasets, metrics, and industrial domains, but its statistics are undermined by internal inconsistencies and m...

  11. Vision-Language Models for Edge Networks: A Comprehensive Survey

    cs.CV 2025-02 reject novelty 2.0 of 10

    A survey of lightweight vision-language models for edge deployment, marred by citation errors, self-citation, and a lack of selection methodology.

Pith tools