Pith. sign in

REVIEW 1 cited by

VoLTA: Vision-Language Transformer with Weakly-Supervised Local-Feature Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.04135 v3 pith:OABC3WZH submitted 2022-10-09 cs.CV cs.LGcs.MM

classification cs.CVcs.LGcs.MM
keywords voltavision-languagealignmentannotationsdownstreamfine-grainedtransformerweakly-supervised
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-language pre-training (VLP) has recently proven highly effective for various uni- and multi-modal downstream applications. However, most existing end-to-end VLP methods use high-resolution image-text box data to perform well on fine-grained region-level tasks, such as object detection, segmentation, and referring expression comprehension. Unfortunately, such high-resolution images with accurate bounding box annotations are expensive to collect and use for supervision at scale. In this work, we propose VoLTA (Vision-Language Transformer with weakly-supervised local-feature Alignment), a new VLP paradigm that only utilizes image-caption data but achieves fine-grained region-level image understanding, eliminating the use of expensive box annotations. VoLTA adopts graph optimal transport-based weakly-supervised alignment on local image patches and text tokens to germinate an explicit, self-normalized, and interpretable low-level matching criterion. In addition, VoLTA pushes multi-modal fusion deep into the uni-modal backbones during pre-training and removes fusion-specific transformer layers, further reducing memory requirements. Extensive experiments on a wide range of vision- and vision-language downstream tasks demonstrate the effectiveness of VoLTA on fine-grained applications without compromising the coarse-grained downstream performance, often outperforming methods using significantly more caption and box annotations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Representations in vision and language converge in a shared, multidimensional space of perceived similarities

    q-bio.NC 2025-07 conditional novelty 6.0 of 10

    Similarity judgments of natural scene images and their sentence captions are aligned with each other, with visual brain responses, and with LLM-trained visual models.

Pith tools