Pith. sign in

REVIEW 2 cited by

Prismer: A Vision-Language Model with Multi-Task Experts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.02506 v3 pith:OABOTDEF submitted 2023-03-04 cs.LG cs.AIcs.CV

classification cs.LGcs.AIcs.CV
keywords prismerexpertstrainingvision-languagemodelmodelsachievesadapt
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent vision-language models have shown impressive multi-modal generation capabilities. However, typically they require training huge models on massive datasets. As a more scalable alternative, we introduce Prismer, a data- and parameter-efficient vision-language model that leverages an ensemble of task-specific experts. Prismer only requires training of a small number of components, with the majority of network weights inherited from multiple readily-available, pre-trained experts, and kept frozen during training. By leveraging experts from a wide range of domains, we show Prismer can efficiently pool this expert knowledge and adapt it to various vision-language reasoning tasks. In our experiments, we show that Prismer achieves fine-tuned and few-shot learning performance which is competitive with current state-of-the-arts, whilst requiring up to two orders of magnitude less training data. Code is available at https://github.com/NVlabs/prismer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 15 citations worldwide. Full citation record

  1. Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Stable Diffusion features, especially when conditioned on the question, improve vision-centric multimodal question answering when fused with CLIP.

  2. VisCon-100K: Leveraging Contextual Web Data for Fine-tuning Vision Language Models

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A web-context-derived dataset and a 'leaky modality mix' of captions with Q&A pairs improve vision-language model fine-tuning on several benchmarks.

Pith tools