Pith. sign in

REVIEW 1 cited by

VLMAE: Vision-Language Masked Autoencoder

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.09374 v1 pith:O5TNFI5A submitted 2022-08-19 cs.CV

classification cs.CV
keywords vlmaeimagevision-languagevisualautoencoderfeaturesimage-textmasked
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image and language modeling is of crucial importance for vision-language pre-training (VLP), which aims to learn multi-modal representations from large-scale paired image-text data. However, we observe that most existing VLP methods focus on modeling the interactions between image and text features while neglecting the information disparity between image and text, thus suffering from focal bias. To address this problem, we propose a vision-language masked autoencoder framework (VLMAE). VLMAE employs visual generative learning, facilitating the model to acquire fine-grained and unbiased features. Unlike the previous works, VLMAE pays attention to almost all critical patches in an image, providing more comprehensive understanding. Extensive experiments demonstrate that VLMAE achieves better performance in various vision-language downstream tasks, including visual question answering, image-text retrieval and visual grounding, even with up to 20% pre-training speedup.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Grad-CAM for Vision Transformers: A Systematic Taxonomy and Audit of Methodological Ambiguity in Explainable AI

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A systematic audit of 175 papers finds that most uses of Grad-CAM on vision transformers omit the implementation choices needed to reproduce the visual explanation, and introduces a taxonomy to name those choices.

Pith tools