Pith. sign in

REVIEW 1 cited by

Attention, Please! PixelSHAP Reveals What Vision-Language Models Actually Focus On

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.06670 v1 pith:GHA3235M submitted 2025-03-09 cs.CV cs.CL

classification cs.CVcs.CL
keywords pixelshapmodelsinterpretabilitymethodsopen-sourceshapley-basedvision-languageability
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Interpretability in Vision-Language Models (VLMs) is crucial for trust, debugging, and decision-making in high-stakes applications. We introduce PixelSHAP, a model-agnostic framework extending Shapley-based analysis to structured visual entities. Unlike previous methods focusing on text prompts, PixelSHAP applies to vision-based reasoning by systematically perturbing image objects and quantifying their influence on a VLM's response. PixelSHAP requires no model internals, operating solely on input-output pairs, making it compatible with open-source and commercial models. It supports diverse embedding-based similarity metrics and scales efficiently using optimization techniques inspired by Shapley-based methods. We validate PixelSHAP in autonomous driving, highlighting its ability to enhance interpretability. Key challenges include segmentation sensitivity and object occlusion. Our open-source implementation facilitates further research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A gradient-attention explainability method produces sequence-level visual and textual saliency maps for free-form answers from large vision-language models, with stronger human-attention alignment and faithfulness tha...

Pith tools