Pith. sign in

REVIEW 2 cited by

Linear Explanations for Individual Neurons

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.06855 v1 pith:WUJKDVUJ submitted 2024-05-10 cs.LG cs.CV

classification cs.LGcs.CV
keywords activationslinearneuronneuronsonlyveryadditionexplanations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years many methods have been developed to understand the internal workings of neural networks, often by describing the function of individual neurons in the model. However, these methods typically only focus on explaining the very highest activations of a neuron. In this paper we show this is not sufficient, and that the highest activation range is only responsible for a very small percentage of the neuron's causal effect. In addition, inputs causing lower activations are often very different and can't be reliably predicted by only looking at high activations. We propose that neurons should instead be understood as a linear combination of concepts, and develop an efficient method for producing these linear explanations. In addition, we show how to automatically evaluate description quality using simulation, i.e. predicting neuron activations on unseen inputs in vision setting.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Data-Efficient Adaptation of LLMs via Attention Head Reweighting

    cs.LG 2026-07 conditional novelty 5.0 of 10

    Learning a single scalar per attention head lets LLMs adapt to few-shot text classification better than LoRA, with 200–1000x fewer trainable parameters.

  2. FastCAV: Efficient Computation of Concept Activation Vectors for Explaining Deep Neural Networks

    cs.LG 2025-05 conditional novelty 3.0 of 10

    Concept activation vectors can be computed as the normalized difference between concept-mean and global-mean activations, giving a 46.4x average speedup over SVM-based CAVs with comparable quality.

Pith tools