Pith. sign in

REVIEW 3 cited by

Multimodal Foundation Models: From Specialists to General-Purpose Assistants

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.10020 v1 pith:7RSYR56W submitted 2023-09-18 cs.CV cs.CL

classification cs.CVcs.CL
keywords modelsmultimodalfoundationvisionassistantsgeneral-purposellmsresearch
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents a comprehensive survey of the taxonomy and evolution of multimodal foundation models that demonstrate vision and vision-language capabilities, focusing on the transition from specialist models to general-purpose assistants. The research landscape encompasses five core topics, categorized into two classes. (i) We start with a survey of well-established research areas: multimodal foundation models pre-trained for specific purposes, including two topics -- methods of learning vision backbones for visual understanding and text-to-image generation. (ii) Then, we present recent advances in exploratory, open research areas: multimodal foundation models that aim to play the role of general-purpose assistants, including three topics -- unified vision models inspired by large language models (LLMs), end-to-end training of multimodal LLMs, and chaining multimodal tools with LLMs. The target audiences of the paper are researchers, graduate students, and professionals in computer vision and vision-language multimodal communities who are eager to learn the basics and recent advances in multimodal foundation models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EVADE-Bench: Multimodal Benchmark for Evaluating and Enhancing Evasive Content Detection

    cs.CL 2025-05 conditional novelty 6.0 of 10

    A new benchmark of 2,833 evasive text samples and 13,961 images shows current LLMs and VLMs frequently miss veiled policy violations in Chinese e-commerce ads.

  2. A Survey on Video Temporal Grounding with Multimodal Large Language Model

    cs.CV 2025-08 unverdicted novelty 3.0 of 10

    A taxonomized review of video temporal grounding with multimodal large language models, covering model roles, training paradigms, feature processing, benchmarks, and open problems.

  3. Vision Generalist Model: A Survey

    cs.CV 2025-06 conditional novelty 3.0 of 10

    A structured review of vision generalist models, classifying them into encoding-based and sequence-to-sequence frameworks and summarizing datasets, benchmarks, techniques, and open problems.

Pith tools