Pith. sign in

REVIEW 1 cited by

Decomposing and Interpreting Image Representations via Text in ViTs Beyond CLIP

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.01583 v2 pith:2TESRFOU submitted 2024-06-03 cs.CV cs.LG

classification cs.CVcs.LG
keywords componentsclipimagefeaturesrepresentationtextvitsbeyond
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent work has explored how individual components of the CLIP-ViT model contribute to the final representation by leveraging the shared image-text representation space of CLIP. These components, such as attention heads and MLPs, have been shown to capture distinct image features like shape, color or texture. However, understanding the role of these components in arbitrary vision transformers (ViTs) is challenging. To this end, we introduce a general framework which can identify the roles of various components in ViTs beyond CLIP. Specifically, we (a) automate the decomposition of the final representation into contributions from different model components, and (b) linearly map these contributions to CLIP space to interpret them via text. Additionally, we introduce a novel scoring function to rank components by their importance with respect to specific features. Applying our framework to various ViT variants (e.g. DeiT, DINO, DINOv2, Swin, MaxViT), we gain insights into the roles of different components concerning particular image features. These insights facilitate applications such as image retrieval using text descriptions or reference images, visualizing token importance heatmaps, and mitigating spurious correlations. We release our code to reproduce the experiments at https://github.com/SriramB-98/vit-decompose

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Quantifying Structure in CLIP Embeddings: A Statistical Framework for Concept Interpretation

    cs.LG 2025-06 reject novelty 6.0 of 10

    A rotation-sensitivity hypothesis test plus Varimax rotation produces sparse concept dictionaries from CLIP embeddings and improves worst-group accuracy after spurious concept removal.

Pith tools