REVIEW 6 cited by
From Colors to Classes: Emergence of Concepts in Vision Transformers
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Vision Transformers (ViTs) are increasingly utilized in various computer vision tasks due to their powerful representation capabilities. However, it remains understudied how ViTs process information layer by layer. Numerous studies have shown that convolutional neural networks (CNNs) extract features of increasing complexity throughout their layers, which is crucial for tasks like domain adaptation and transfer learning. ViTs, lacking the same inductive biases as CNNs, can potentially learn global dependencies from the first layers due to their attention mechanisms. Given the increasing importance of ViTs in computer vision, there is a need to improve the layer-wise understanding of ViTs. In this work, we present a novel, layer-wise analysis of concepts encoded in state-of-the-art ViTs using neuron labeling. Our findings reveal that ViTs encode concepts with increasing complexity throughout the network. Early layers primarily encode basic features such as colors and textures, while later layers represent more specific classes, including objects and animals. As the complexity of encoded concepts increases, the number of concepts represented in each layer also rises, reflecting a more diverse and specific set of features. Additionally, different pretraining strategies influence the quantity and category of encoded concepts, with finetuning to specific downstream tasks generally reducing the number of encoded concepts and shifting the concepts to more relevant categories.
Forward citations
Cited by 6 Pith papers
-
MIMIC: Multimodal Inversion for Model Interpretation and Conceptualization
MIMIC is a new inversion framework that recovers visual concepts from VLM internal states using joint inversion, feature alignment, and three regularizers.
-
The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail
Reliable concept presence in transformers is concentrated in the extreme high-activation tail of in-concept tokens; thresholding that tail improves concept detection and localization.
-
Neutralizing Token Aggregation via Information Augmentation for Efficient Test-Time Adaptation
NAVIA augments the [CLS] token with adaptive biases in shallow ViT layers so entropy-minimizing test-time adaptation can recover information lost to token aggregation, reporting over 2.5% accuracy gains and over 20% l...
-
Granular Concept Circuits: Toward a Fine-Grained Circuit Discovery for Concept Representations
GCC discovers multiple concept-specific neuron circuits per query by combining first-order ablation sensitivity with top-k activation overlap.
-
Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models
Orthogonality-regularizing a language model's sparse-autoencoder features modestly improves the model's ability to swap a named entity during generation, without hurting math-reasoning accuracy.
-
Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models
Enforcing near-orthogonality on sparse-autoencoder features in a fine-tuned language model improves the isolation of concept interventions while keeping math performance roughly unchanged.
Discussion (0). Sign in to comment.