Pith. sign in

REVIEW 18 cited by

Knowledge Neurons in Pretrained Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2104.08696 v2 pith:IPF5UAAN submitted 2021-04-18 cs.CL

classification cs.CL
keywords knowledgeneuronspretrainedfactualtransformersfactstudiesactivation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale pretrained language models are surprisingly good at recalling factual knowledge presented in the training corpus. In this paper, we present preliminary studies on how factual knowledge is stored in pretrained Transformers by introducing the concept of knowledge neurons. Specifically, we examine the fill-in-the-blank cloze task for BERT. Given a relational fact, we propose a knowledge attribution method to identify the neurons that express the fact. We find that the activation of such knowledge neurons is positively correlated to the expression of their corresponding facts. In our case studies, we attempt to leverage knowledge neurons to edit (such as update, and erase) specific factual knowledge without fine-tuning. Our results shed light on understanding the storage of knowledge within pretrained Transformers. The code is available at https://github.com/Hunter-DDM/knowledge-neurons.

Discussion (0). Sign in to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 42 citations worldwide. Full citation record

  1. Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens

    cs.CV 2025-09 conditional novelty 7.0 of 10

    Feed-forward neurons in LVLMs encode whether a text token is visually grounded, and a detector built on these neurons can reduce hallucination by overriding or replacing ungrounded tokens.

  2. Fairness Pruning: Locating Demographic Bias in GLU-MLP Layers via Differential Activations

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Unsigned differential activations locate a few GLU-MLP neurons whose zeroing surgically destabilizes demographic bias while retaining ~99.5% of measured capabilities.

  3. Granular Concept Circuits: Toward a Fine-Grained Circuit Discovery for Concept Representations

    cs.CV 2025-08 conditional novelty 6.0 of 10

    GCC discovers multiple concept-specific neuron circuits per query by combining first-order ablation sensitivity with top-k activation overlap.

  4. What Should LLMs Forget? Quantifying Personal Data in LLMs for Right-to-Be-Forgotten Requests

    cs.CL 2025-07 conditional novelty 6.0 of 10

    WikiMem, a Wikidata-derived canary dataset and a calibrated NLL-ranking metric, identifies which human-fact associations an LLM has memorized, with higher rates for famous people and larger models.

  5. Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Shortcut neuron patching suppresses benchmark-contamination shortcuts in LLMs and yields evaluation scores that strongly correlate with the external MixEval benchmark.

  6. TRACE for Tracking the Emergence of Semantic Representations in Transformers

    cs.CL 2025-05 reject novelty 6.0 of 10

    Using Hessian curvature, intrinsic dimensionality, and linguistic probes on a synthetic frame-semantic corpus, the paper claims a coordinated intersection-based phase transition in small transformers, though the marke...

  7. Break Through the Compression Bottleneck: From Theory to Practice

    cs.CL 2026-05 reject novelty 5.0 of 10

    The paper asserts a first proof that low-rank decomposition and quantization are non-orthogonal tools for LLM compression, recommends low-rank-first ordering, and adds a diagonal scaling fix (DAM) that reduces the com...

  8. QF: Quick Feedforward AI Model Training without Gradient Back Propagation

    cs.LG 2025-07 reject novelty 5.0 of 10

    QF Learning updates transformer weights with a closed-form, backprop-free formula so that a model can recall an injected fact from memory after a single instructional example, but the reported evidence is only qualitative.

  9. SplitLoRA: Balancing Stability and Plasticity in Continual Learning Through Gradient Space Splitting

    cs.LG 2025-05 reject novelty 5.0 of 10

    SplitLoRA picks the LoRA update subspace size from previous-task gradient singular values using a hyperparameter alpha, and freezes the projection to keep updates in that subspace.

  10. A Graph Perspective to Probe Structural Patterns of Knowledge in Large Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    LLM knowledge, measured by self-reported true/false checks on knowledge-graph triplets, shows homophily and degree correlations that a graph neural network exploits to select more effective fine-tuning data.

  11. Activation Control for Efficiently Eliciting Long Chain-of-thought Ability of Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Training-free amplification of selected last-layer activations, combined with 'wait' token insertion, elicits long chain-of-thought reasoning in base LLMs and improves accuracy on math and science benchmarks.

  12. Locate-then-Merge: Neuron-Level Parameter Fusion for Mitigating Catastrophic Forgetting in Multimodal LLMs

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Neuron-Fusion selectively restores large-change neurons from a fine-tuned multimodal model and suppresses small changes, improving language retention with modest visual trade-offs.

  13. Towards a Science of Causal Interpretability in Deep Learning for Software Engineering

    cs.SE 2025-05 conditional novelty 5.0 of 10

    The dissertation presents docode, a causal interpretability method for neural code models, and uses a case study to show that some correlations between code properties and model performance are confounded rather than causal.

  14. Context-Adaptive Inference: A Unified Statistical and Foundation-Model View

    stat.ML 2026-07 conditional novelty 4.0 of 10

    Under linear, squared-loss assumptions, explicit context adaptation and in-context learning both reduce to kernel ridge regression on joint input-context features.

  15. Rhetorical Text-to-Image Generation via Two-layer Diffusion Policy Optimization

    cs.CV 2025-05 reject novelty 4.0 of 10

    Rhet2Pix combines staged LLM prompt decomposition with a discounted PPO fine-tuning scheme for Stable Diffusion, claiming strong rhetorical text-to-image generation, but the quantitative evidence is circular and undefined.

  16. Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Activation-frequency analysis identifies sparse units in LLMs that respond to instructions; same-category instructions share more of these units than different-category ones, and fine-tuning measurably changes the sets.

  17. Neural Incompatibility: The Unbridgeable Gap of Cross-Scale Parametric Knowledge Transfer in Large Language Models

    cs.CL 2025-05 conditional novelty 4.0 of 10

    Directly transferring parameters between differently-sized language models is unreliable; the paper proposes a pre-alignment method (LaTen) and explains the failure via 'Neural Incompatibility'.

  18. A Survey on Proactive Defense Strategies Against Misinformation in Large Language Models

    cs.IR 2025-07 reject novelty 3.0 of 10

    A survey claims proactive defenses against LLM misinformation outperform post-hoc detection by up to 63%, but no meta-analysis details are provided to support the claim.

Pith tools