Pith. sign in

REVIEW 3 cited by

Interpreting Neurons in Deep Vision Networks with Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.13771 v2 pith:JMYSWKYP submitted 2024-03-20 cs.CV cs.LG

classification cs.CVcs.LG
keywords modelsbestdatadeepdescribe-and-dissectdescriptionslanguagemethod
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

In this paper, we propose Describe-and-Dissect (DnD), a novel method to describe the roles of hidden neurons in vision networks. DnD utilizes recent advancements in multimodal deep learning to produce complex natural language descriptions, without the need for labeled training data or a predefined set of concepts to choose from. Additionally, DnD is training-free, meaning we don't train any new models and can easily leverage more capable general purpose models in the future. We have conducted extensive qualitative and quantitative analysis to show that DnD outperforms prior work by providing higher quality neuron descriptions. Specifically, our method on average provides the highest quality labels and is more than 2$\times$ as likely to be selected as the best explanation for a neuron than the best baseline. Finally, we present a use case providing critical insights into land cover prediction models for sustainability applications. Our code and data are available at https://github.com/Trustworthy-ML-Lab/Describe-and-Dissect.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NAT: Learning to Attack Neurons for Enhanced Adversarial Transferability

    cs.CV 2025-08 conditional novelty 6.0 of 10

    NAT trains per-neuron adversarial generators that each disrupt one mid-layer neuron, improving cross-model and cross-domain attack transferability over embedding-level baselines.

  2. Mechanistic understanding and validation of large AI models with SemanticLens

    cs.LG 2025-01 conditional novelty 6.0 of 10

    SemanticLens maps each neuron of a vision model to a CLIP-space vector, enabling text-based search, labelling, audit, and interpretability scoring of model internals.

  3. Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions

    cs.CY 2026-02 unverdicted novelty 4.0 of 10

    Current XAI methods for DNNs and LLMs rest on paradoxes and false assumptions that demand a paradigm shift to verification protocols, scientific foundations, context-aware design, and faithful model analysis rather th...

Pith tools