Pith. sign in

REVIEW 2 cited by

Vision-Language Pre-training: Basics, Recent Advances, and Future Trends

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.09263 v1 pith:2F7OTUZX submitted 2022-10-17 cs.CV cs.CL

classification cs.CVcs.CL
keywords tasksansweringbeencaptioningcategorycomputerdiscussimage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

This paper surveys vision-language pre-training (VLP) methods for multimodal intelligence that have been developed in the last few years. We group these approaches into three categories: ($i$) VLP for image-text tasks, such as image captioning, image-text retrieval, visual question answering, and visual grounding; ($ii$) VLP for core computer vision tasks, such as (open-set) image classification, object detection, and segmentation; and ($iii$) VLP for video-text tasks, such as video captioning, video-text retrieval, and video question answering. For each category, we present a comprehensive review of state-of-the-art methods, and discuss the progress that has been made and challenges still being faced, using specific systems and models as case studies. In addition, for each category, we discuss advanced topics being actively explored in the research community, such as big foundation models, unified modeling, in-context few-shot learning, knowledge, robustness, and computer vision in the wild, to name a few.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CLIP-UP: A Simple and Efficient Mixture-of-Experts CLIP Training Recipe with Sparse Upcycling

    cs.CV 2025-02 conditional novelty 5.0 of 10

    CLIP-UP converts a pre-trained dense CLIP into an MoE model and improves zero-shot text-image retrieval beyond dense baselines at lower inference cost.

  2. Vision-Language Models for Edge Networks: A Comprehensive Survey

    cs.CV 2025-02 reject novelty 2.0 of 10

    A survey of lightweight vision-language models for edge deployment, marred by citation errors, self-citation, and a lack of selection methodology.

Pith tools