Pith. sign in

REVIEW 4 cited by

What Do VLMs NOTICE? A Mechanistic Interpretability Pipeline for Gaussian-Noise-free Text-Image Corruption and Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.16320 v3 pith:MDLVHLEX submitted 2024-06-24 cs.CL

classification cs.CL
keywords corruptionmodelsnoticevlmscross-attentiondecision-makingevaluationheads
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Vision-Language Models (VLMs) have gained community-spanning prominence due to their ability to integrate visual and textual inputs to perform complex tasks. Despite their success, the internal decision-making processes of these models remain opaque, posing challenges in high-stakes applications. To address this, we introduce NOTICE, the first Noise-free Text-Image Corruption and Evaluation pipeline for mechanistic interpretability in VLMs. NOTICE incorporates a Semantic Minimal Pairs (SMP) framework for image corruption and Symmetric Token Replacement (STR) for text. This approach enables semantically meaningful causal mediation analysis for both modalities, providing a robust method for analyzing multimodal integration within models like BLIP. Our experiments on the SVO-Probes, MIT-States, and Facial Expression Recognition datasets reveal crucial insights into VLM decision-making, identifying the significant role of middle-layer cross-attention heads. Further, we uncover a set of ``universal cross-attention heads'' that consistently contribute across tasks and modalities, each performing distinct functions such as implicit image segmentation, object inhibition, and outlier inhibition. This work paves the way for more transparent and interpretable multimodal systems.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A new taxonomy and 1,500-item dataset, ENTRAP-VL, lets researchers measure whether vision-language models are entrained by textual and visual context separately.

  2. What Do They See? Interpreting Complex Road Scenarios Through the Eyes of Vision-Language-Action Models for Safe and Trustworthy Autonomous Vehicle Learning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Erasing objects from front-camera images shows Alpamayo 1's trajectories depend most on large vehicles, pedestrians, and traffic lights, but attributions are seed-unstable and some effects reach the output without tou...

  3. How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA

    cs.CV 2026-07 reject novelty 6.0 of 10

    The paper proposes four operation-level VLM failure modes and a pathway dissociation, but the dissociation is not supported by the paper's own intervention statistics.

  4. CAST: Mitigating Object Hallucination in Large Vision-Language Models via Caption-Guided Visual Attention Steering

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    CAST reduces object hallucination in LVLMs by 6.03% on average across five models and five benchmarks by identifying caption-sensitive attention heads and applying optimized steering directions to their outputs, with ...

Pith tools