Pith. sign in

REVIEW 2 cited by

VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.02358 v2 pith:45XKWBBN submitted 2021-11-03 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords vlmoencodervision-languageimage-textpretraineddualfusionmixture-of-modality-experts
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present a unified Vision-Language pretrained Model (VLMo) that jointly learns a dual encoder and a fusion encoder with a modular Transformer network. Specifically, we introduce Mixture-of-Modality-Experts (MoME) Transformer, where each block contains a pool of modality-specific experts and a shared self-attention layer. Because of the modeling flexibility of MoME, pretrained VLMo can be fine-tuned as a fusion encoder for vision-language classification tasks, or used as a dual encoder for efficient image-text retrieval. Moreover, we propose a stagewise pre-training strategy, which effectively leverages large-scale image-only and text-only data besides image-text pairs. Experimental results show that VLMo achieves state-of-the-art results on various vision-language tasks, including VQA, NLVR2 and image-text retrieval. The code and pretrained models are available at https://aka.ms/vlmo.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FREE: Fast and Robust Vision Language Models with Early Exits

    cs.LG 2025-06 conditional novelty 6.0 of 10

    An adversarial early-exit method for frozen-backbone vision language models that reuses the final classifier and reports 1.5x inference speedup with comparable accuracy.

  2. Unsupervised Transcript-assisted Video Summarization and Highlight Detection

    cs.CV 2025-05 conditional novelty 4.0 of 10

    Combining video frames with transcript text in an RL-trained summarizer improves rank-based highlight metrics and summary fidelity on Mr. HiSum, while slightly hurting F1 and top-5% highlight detection.

Pith tools