Pith. sign in

REVIEW 8 cited by

A Survey of Vision-Language Pre-Trained Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.10936 v2 pith:LIBXGDPW submitted 2022-02-18 cs.CV cs.CLcs.LG

classification cs.CVcs.CLcs.LG
keywords modelspre-trainedpre-trainingdownstreamintroducelearningmainstreamrecent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As transformer evolves, pre-trained models have advanced at a breakneck pace in recent years. They have dominated the mainstream techniques in natural language processing (NLP) and computer vision (CV). How to adapt pre-training to the field of Vision-and-Language (V-L) learning and improve downstream task performance becomes a focus of multimodal learning. In this paper, we review the recent progress in Vision-Language Pre-Trained Models (VL-PTMs). As the core content, we first briefly introduce several ways to encode raw images and texts to single-modal embeddings before pre-training. Then, we dive into the mainstream architectures of VL-PTMs in modeling the interaction between text and image representations. We further present widely-used pre-training tasks, and then we introduce some common downstream tasks. We finally conclude this paper and present some promising research directions. Our survey aims to provide researchers with synthesis and pointer to related research.

Discussion (0). Sign in to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Continual Learning for VLMs: A Survey and Taxonomy Beyond Forgetting

    cs.CV 2025-08 unverdicted novelty 7.0 of 10

    The paper offers a comprehensive survey and proposes a new taxonomy for continual learning strategies in VLMs and MLLMs to combat catastrophic forgetting beyond traditional methods.

  2. TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    TouchThinker introduces a 1M-scale multi-source tactile dataset and action-aware modeling to scale commonsense reasoning from tactile observations, reporting competitive performance on new and existing benchmarks.

  3. LEGO Co-builder: Exploring Fine-Grained Vision-Language Modeling for Multimodal LEGO Assembly Assistants

    cs.AI 2025-07 conditional novelty 6.0 of 10

    LEGO-VLM benchmark shows VLMs fail at fine-grained LEGO assembly state detection, but model-generated ground truth and a missing trivial baseline weaken the claim.

  4. MrM: Black-Box Membership Inference Attacks against Multimodal RAG Systems

    cs.CV 2025-06 conditional novelty 6.0 of 10

    MrM is a black-box membership inference attack on multimodal RAG systems that masks key objects in a target image and uses the system's ability to reconstruct them as a membership signal.

  5. Mitigating Object Hallucination via Robust Local Perception Search

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A training-free decoding method that uses an MLLM's own local object descriptions as a reward prior, combined with CLIP similarity, to cut object hallucination, especially under adversarial image noise.

  6. Innocence in the Crossfire: Roles of Skip Connections in Jailbreaking Visual Language Models

    cs.CL 2025-07 conditional novelty 5.0 of 10

    The paper reports higher harmful-output rates in three open-source VLMs from detailed image descriptions, in-context examples, and positive openings, and from a skip connection between internal layers, with memes riva...

  7. CF-VLM:CounterFactual Vision-Language Fine-tuning

    cs.LG 2025-06 conditional novelty 5.0 of 10

    CF-VLM fine-tunes VLMs on counterfactual image-text pairs with three objectives, reporting gains on compositional reasoning benchmarks and modest hallucination reductions.

  8. Generalizing vision-language models to novel domains: A comprehensive survey

    cs.CV 2025-06 conditional novelty 3.0 of 10

    A survey of VLM generalization literature organized by transferred module, with benchmark tables and a review of multimodal LLMs.

Pith tools