REVIEW 8 cited by
Interpreting CLIP's Image Representation via Text-Based Decomposition
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We investigate the CLIP image encoder by analyzing how individual model components affect the final representation. We decompose the image representation as a sum across individual image patches, model layers, and attention heads, and use CLIP's text representation to interpret the summands. Interpreting the attention heads, we characterize each head's role by automatically finding text representations that span its output space, which reveals property-specific roles for many heads (e.g. location or shape). Next, interpreting the image patches, we uncover an emergent spatial localization within CLIP. Finally, we use this understanding to remove spurious features from CLIP and to create a strong zero-shot image segmenter. Our results indicate that a scalable understanding of transformer models is attainable and can be used to repair and improve models.
Forward citations
Cited by 8 Pith papers
-
ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features
ConceptAttention shows that linear projections in the output space of DiT attention layers yield sharper concept-localizing saliency maps than cross-attention maps, reaching state-of-the-art zero-shot segmentation.
-
BrainExplore: Large-Scale Discovery of Interpretable Visual Representations in the Human Brain
A new automated pipeline decomposes fMRI activity into components and labels them with visual concepts, claiming thousands of interpretable patterns across the human visual cortex.
-
The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail
Reliable concept presence in transformers is concentrated in the extreme high-activation tail of in-concept tokens; thresholding that tail improves concept detection and localization.
-
How Visual Representations Map to Language Feature Space in Multimodal LLMs
Visual tokens in a fully frozen-backbone VLM with a linear adapter only become well-represented by the LLM's sparse autoencoder features in middle-to-late layers, converging around layer 18.
-
Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads
Using wrong/correct hard-sample head comparisons, LTC finds spurious CLIP attention heads and corrects them to raise worst-group accuracy on biased benchmarks.
-
Understanding Design Fixation in Generative AI
Generative AI models exhibit a design fixation phenomenon that limits the diversity and originality of their design outputs, according to a small lab study and a proposed theoretical framework.
-
MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence
MEDIC-AD adds anomaly-aware and difference tokens to a medical VLM, claiming SOTA lesion detection, temporal tracking, and visual grounding; the zero-shot claim is undermined by likely train/test overlap.
-
Model Science: getting serious about verification, explanation and control of AI systems
Proposes 'Model Science' as a model-centric paradigm for AI with four pillars: verification, explanation, control, and interface.
Discussion (0). Continue with ORCID to comment.