Pith. sign in

REVIEW 1 cited by

Interpreting and Controlling Vision Foundation Models via Text Explanations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.10591 v1 pith:L6EMPCJE submitted 2023-10-16 cs.CV

classification cs.CV
keywords modelvisionframeworkmodelsbehaviorscontrollingfoundationinterpreting
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale pre-trained vision foundation models, such as CLIP, have become de facto backbones for various vision tasks. However, due to their black-box nature, understanding the underlying rules behind these models' predictions and controlling model behaviors have remained open challenges. We present a framework for interpreting vision transformer's latent tokens with natural language. Given a latent token, our framework retains its semantic information to the final layer using transformer's local operations and retrieves the closest text for explanation. Our approach enables understanding of model visual reasoning procedure without needing additional model training or data collection. Based on the obtained interpretations, our framework allows for model editing that controls model reasoning behaviors and improves model robustness against biases and spurious correlations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. How Visual Representations Map to Language Feature Space in Multimodal LLMs

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Visual tokens in a fully frozen-backbone VLM with a linear adapter only become well-represented by the LLM's sparse autoencoder features in middle-to-late layers, converging around layer 18.

Pith tools