Pith. sign in

REVIEW 7 cited by

Revisiting the Role of Language Priors in Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.01879 v4 pith:GVUBWLVX submitted 2023-06-02 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords vision-languageaccuracybenchmarksgenerativeimagelanguageprobabilisticretrieval
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Vision-language models (VLMs) are impactful in part because they can be applied to a variety of visual understanding tasks in a zero-shot fashion, without any fine-tuning. We study $\textit{generative VLMs}$ that are trained for next-word generation given an image. We explore their zero-shot performance on the illustrative task of image-text retrieval across 8 popular vision-language benchmarks. Our first observation is that they can be repurposed for discriminative tasks (such as image-text retrieval) by simply computing the match score of generating a particular text string given an image. We call this probabilistic score the $\textit{Visual Generative Pre-Training Score}$ (VisualGPTScore). While the VisualGPTScore produces near-perfect accuracy on some retrieval benchmarks, it yields poor accuracy on others. We analyze this behavior through a probabilistic lens, pointing out that some benchmarks inadvertently capture unnatural language distributions by creating adversarial but unlikely text captions. In fact, we demonstrate that even a "blind" language model that ignores any image evidence can sometimes outperform all prior art, reminiscent of similar challenges faced by the visual-question answering (VQA) community many years ago. We derive a probabilistic post-processing scheme that controls for the amount of linguistic bias in generative VLMs at test time without having to retrain or fine-tune the model. We show that the VisualGPTScore, when appropriately debiased, is a strong zero-shot baseline for vision-language understanding, oftentimes producing state-of-the-art accuracy.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ARGUS: Hallucination and Omission Evaluation in Video-LLMs

    cs.CV 2025-06 conditional novelty 7.0 of 10

    ARGUS measures hallucination and omission in free-form video captions using LLM-based entailment and temporal alignment, finding that even the best video-LLM still produces roughly 40% hallucinated content.

  2. Prior Bias in Vision Language Models on UML Diagram Interpretation

    cs.CV 2026-07 conditional novelty 6.5 of 10

    Reversing only the UML relation arrow while keeping class names and layout fixed cuts open-source VLM relation accuracy by about 33%, revealing prior-over-vision bias.

  3. Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A recalibrated GRPO reinforcement learning method lets multimodal LLMs say 'None' for nonexistent referring expressions without sacrificing localization accuracy on objects that do exist.

  4. Scalable Visual Pretraining for Language Intelligence

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    Visual pretraining on rendered document pages beats text-only continued pretraining on scientific reasoning benchmarks across multiple model backbones, using about a quarter of the tokens.

  5. TokenSwap: Backdoor Attack on the Compositional Understanding of Large Vision-Language Models

    cs.CV 2025-09 conditional novelty 6.0 of 10

    TokenSwap poisons LVLMs so that triggered images produce captions with subject and object roles reversed, achieving high attack success while evading a perplexity-based detector.

  6. Trade-offs in Image Generation: How Do Different Dimensions Interact?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new benchmark and VLM-as-judge metric map trade-offs among ten image-generation dimensions across 14 models, with a visualization called DTM.

  7. Decomposing Complex Visual Comprehension into Atomic Visual Skills for Vision Language Models

    cs.CV 2025-05 conditional novelty 5.0 of 10

    VLMs score far below adult humans on a new 13,188-question benchmark of 36 atomic 2D geometry perception skills.

Pith tools