Pith. sign in

Paper Citation Record · LEDGER

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

As of 9 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 9 inbound Pith citation observations for arXiv:2507.02559.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02559 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:31:18.269637Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:11:43.051506Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T17:35:51.305374Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c869a743-df52-4b73-b291-c0d95f896c6b · outbound

This paper cites an unresolved cited work.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.056089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.056089Z digest=sha256:f0722eb3555b017cf34a1771bf03b2b21da3b6f82725b8e17277bf85cce23250

Observation c0c8e58f-c08b-40df-acbc-810dffbea027 · outbound

This paper cites an unresolved cited work.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.172686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.172686Z digest=sha256:057725cee871096e8d4dc9be4e783a17c951d84ccff3e6a391c9ed7cb7b44686

Observation e041b568-94b6-450a-a69f-aece269aa919 · outbound

This paper cites Why do LLMs attend to the first token?.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Why do LLMs attend to the first token?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.257797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.257797Z digest=sha256:63b1b79460e84513a3a0aa5d603d9b52e1609c40656ad89163d7e696bef45854

Observation 11fdb6d2-f6eb-4844-bea2-6740f89faba2 · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Towards monosemanticity: Decomposing language models with dictionary learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.344778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.344778Z digest=sha256:0f81ea7974b9f78a0965502446eb5814aa96bb56f5f18cc89d714b39ed4126d7

Observation dde8764f-8c93-4737-b02f-420c9baa97b3 · outbound

This paper cites The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability The Local Interaction Basis: Identifying Computationally-Relevant and Sparsely Interacting Features in Neural Networks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.407102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.407102Z digest=sha256:d7462657d972c27cf1d27b90c73f63121d1a152bc5c5a04049e26293e8ec4ca6

Observation 42a3c983-8edd-4529-9f24-b064e7bff7fc · outbound

This paper cites A mathematical framework for transformer circuits.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability A mathematical framework for transformer circuits

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.553041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.553041Z digest=sha256:d2a49fbc35283ddf31ba6b83b3569231540dff3a8409df9913e4715781a4599e

Observation e68b7bf3-5b83-4e2e-bd7f-67b328e240c4 · outbound

This paper cites Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Jacobian Sparse Autoencoders: Sparsify Computations, Not Just Activations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.674321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.674321Z digest=sha256:17dc26a8a2a5f872dbde4e0044efd59cb5ba935b91424bf00ec222c67dbf2bff

Observation d20ea447-be10-4c78-90b7-ed412a0a938a · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.779455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.779455Z digest=sha256:c723339ce832b02203e0cf594778f2da7638a06f50908162e1cb59be955a9ad0

Observation 798e706f-9568-4c43-9375-68cdc5b1949b · outbound

This paper cites Finding alignments between interpretable causal variables and distributed neural representations.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Finding alignments between interpretable causal variables and distributed neural representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.884759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.884759Z digest=sha256:2b37584136f914bdf22ae8a519560f204c2f3567617a08309978fec8c4af19c6

Observation e73c837d-6e25-471b-a3a7-fab05156562f · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Gemini: A Family of Highly Capable Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:15.963355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:15.963355Z digest=sha256:844547c7c3adc2de35bd9faa7b64168d7440f8bd2a043a5119923275f20bc2b5

Observation fb4e697c-81f2-4f85-a760-a354b24aa4d9 · outbound

This paper cites Openwebtext corpus.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Openwebtext corpus

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.060049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.060049Z digest=sha256:2f42002de7f93b162374ca19d995d9f3e9ffa3244312b3ac4e72438ca93bc97b

Observation 111d19bb-32ae-4288-b298-63ef7fc93219 · outbound

This paper cites Geometric interpretation of layer normalization and a comparative analysis with rmsnorm, 2025.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Geometric interpretation of layer normalization and a comparative analysis with rmsnorm, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:21.767788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.177715Z digest=sha256:6cb42f95967179d98eac2a3dbba107a030387a46153f2053f488639679d050b0

Observation 7239c004-3271-4d79-b952-90e028f687bd · outbound

This paper cites Universal neurons in gpt2 language models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Universal neurons in gpt2 language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:21.597286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.231561Z digest=sha256:ad43178303b10d92f8ab4b2cb669727012bb20444d66f2a8ec16f9a984b5ea48

Observation 99434904-f61f-4b34-a984-1539183175a3 · outbound

This paper cites You can remove gpt2's LayerNorm by fine-tuning.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability You can remove gpt2's LayerNorm by fine-tuning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:21.426689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.288441Z digest=sha256:dd800e95c27b7ee6aaead94ccee008b3ac7eb847b9c8cd25dc8fb02d20829437

Observation fe9d0fef-a56e-4bdb-be57-c2ff16090900 · outbound

This paper cites How to use and interpret activation patching.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability How to use and interpret activation patching

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.371013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.371013Z digest=sha256:7bae03134e9e1efc9589f9cb631ded7647cb7a55f7cf2c8757e3518010d5f265

Observation b1b3b626-81be-4d72-90fb-0700bf0324e7 · outbound

This paper cites Batch normalization: Accelerating deep network training by reducing internal covariate shift.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Batch normalization: Accelerating deep network training by reducing internal covariate shift

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:21.086538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.500310Z digest=sha256:02b97972a4032a389235e430bdeb27df028f5efb1c746c057bfc15639bc43729

Observation d04f406e-d069-4c5b-9e81-7b1851e40253 · outbound

This paper cites Visit: Visualizing and interpreting the semantic information flow of transformers.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Visit: Visualizing and interpreting the semantic information flow of transformers

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.802792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.581740Z digest=sha256:2969547e0caae713173dd2d0062ab3abd83f00a59dd866efc76c4aa1f3f2970e

Observation a35f9380-5228-4bcf-a980-eda9ca20e0e7 · outbound

This paper cites Sparse autoencoders work on attention layer outputs, Jan 2024.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Sparse autoencoders work on attention layer outputs, Jan 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.597517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:31:16.661983Z digest=sha256:d2af3c7cb078582f0babb109925d69f2ed8e7e72f0b5d106883fc03cc1fdcb73

Observation 2c399d03-4f48-4f46-bda9-872795db513b · outbound

This paper cites an unresolved cited work.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.726988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.726988Z digest=sha256:f6942d795e73fde6fb809aef78c9308d8926499c2f44dced72cc5243146477f6

Observation 546eccdf-40c9-4f50-8908-d8228a2951a3 · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.822219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.822219Z digest=sha256:6c9a11eb978e7987e5705d09a9f8ca82f254d0383e77c1513d24b5d52af55c1f

Observation d07e6a75-ee99-41cb-86a3-0bb1d54231b9 · outbound

This paper cites Copy Suppression: Comprehensively Understanding an Attention Head.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Copy Suppression: Comprehensively Understanding an Attention Head

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:16.921232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:16.921232Z digest=sha256:18794b8600be639a5d14a537dcf5eb042ebffaf429c60add2333e5b385102cf7

Observation b3392b0e-1386-4c5d-8410-29acc96e6ab1 · outbound

This paper cites Locating and Editing Factual Associations in GPT.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Locating and Editing Factual Associations in GPT

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.012912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.012912Z digest=sha256:f07ad76dbef26595c65168ab64bfc6bfc53a5eaedec680d196720236a8ed186f

Observation 9be48507-3147-49c3-8d30-9f26f23fa5fe · outbound

This paper cites Tinymodel: A tinystories lm with saes and transcoders, 2024.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Tinymodel: A tinystories lm with saes and transcoders, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.467692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.087598Z digest=sha256:1e240ba15ed6d3be7895be4ce862b5c6141d238cbf3f5a9cf6821f6d78319842

Observation aa89992e-cd0f-4ba6-8fe5-32689f46f113 · outbound

This paper cites Attribution patching: Activation patching at industrial scale, Mar 2023 a.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Attribution patching: Activation patching at industrial scale, Mar 2023 a

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.385237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.142216Z digest=sha256:37716e3632ca46e5abede876aed02b87734f2e0ba621ba942c42010ebd1d913b

Observation 4c84a2a1-d6f1-4065-82e7-f136d1c90939 · outbound

This paper cites Exploratory analysis demo (transformerlens).

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Exploratory analysis demo (transformerlens)

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.225301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.226190Z digest=sha256:65e0e56ab92c3a163abdada87dc0ddd5252660435d4341863503bde3c83be6e6

Observation 1dd69c8d-38aa-4cca-8965-b2d0889424cd · outbound

This paper cites Transformerlens.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Transformerlens

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.293516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.293516Z digest=sha256:7a847e372571d49041b003706215e545efb155725891045a73d231cf3aa3fb86

Observation f971a940-3747-46ba-bae3-76ae16767369 · outbound

This paper cites interpreting gpt: the logit lens, Aug 2020.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability interpreting gpt: the logit lens, Aug 2020

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:20.082358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.362337Z digest=sha256:9534d4366b68e8670e1b827f1329c96c1b85f10aa516e9ae9893d850ad8dc1f9

Observation e749d544-cfd0-4e1b-9617-f80df5827e35 · outbound

This paper cites an unresolved cited work.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.418514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.418514Z digest=sha256:d04aafbee00891183e4f2edb0c35e30ca4052dc451f69dbb1bfed0379aa18cd6

Observation 4ddc0416-7fcc-431f-8e6a-17f5efd33b43 · outbound

This paper cites Direct and indirect effects.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Direct and indirect effects

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:19.538617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.474033Z digest=sha256:f407068649c02ad7cd452a0f0ab68cec1a79996c277ffb71d1cf65abdbfcb338

Observation 9a41a79c-d3f2-494b-bef6-e251be332513 · outbound

This paper cites Confidence Regulation Neurons in Language Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Confidence Regulation Neurons in Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.531668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.531668Z digest=sha256:4909349498e9651a1fdc9d48759c62397b61b0e1043df112e29a669000fb7590

Observation f1e544b0-9800-4f08-b4d3-7f091708f47a · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability LLaMA: Open and Efficient Foundation Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.573752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.573752Z digest=sha256:e647720714f5a31827dd8a52432331790ef4147cb780d04a05324aecabedc5cd

Observation 9a14a053-66b1-48d6-9c44-fe69d155270f · outbound

This paper cites Attention is all you need.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Attention is all you need

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.639147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.639147Z digest=sha256:bab739a40f1a89c13af4c1c820f971d68b4c78fccbdd26c7a24901acaf2727fb

Observation 2c55ad8d-8a26-4887-b6e2-0b7c0c681830 · outbound

This paper cites Understanding the failure of batch normalization for transformers in nlp.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Understanding the failure of batch normalization for transformers in nlp

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:19.058706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.702651Z digest=sha256:b68e02eed94ebb35c64ff5b929b68f52f9d8294776a1b693805e7558e2dde406

Observation a1e61622-ebda-4274-882d-000351d125d4 · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.798498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.798498Z digest=sha256:566bd2fdff6564f20120f496f27e9a5df6da1613b409c5007888af3db5f5c9a4

Observation 51304b89-5000-4a5d-90fe-45a9839812db · outbound

This paper cites Re-examining layernorm.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Re-examining layernorm

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:18.869243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:31:17.865032Z digest=sha256:b8ab94cf1341ef1b3baa65bfbfb5159fcd5f29e1a5d07c78b1916a048faa179e

Observation 36a28ddc-1ade-4255-aca7-410308e28a62 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Efficient Streaming Language Models with Attention Sinks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:17.953864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:17.953864Z digest=sha256:f17e992df42625145bfd88216dd46131757d0ac1eb053b4803804a78db449eba

Observation 0e012156-7341-48ac-9b01-fb96607b4c76 · outbound

This paper cites Interpreting the Repeated Token Phenomenon in Large Language Models.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Interpreting the Repeated Token Phenomenon in Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:18.044410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:18.044410Z digest=sha256:d70c3dcfc2433db4fe3d2a93fbdadefc382257679463adfc04096a5cec42332e

Observation 44f04f6d-ef34-414e-be54-aa848e4d1365 · outbound

This paper cites Root mean square layer normalization, 2019.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Root mean square layer normalization, 2019

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:18.138243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:18.138243Z digest=sha256:939be63f8c488e05a6d5ca58a05b2c2b3fc575e8f13fff165bbc572e52678754

Observation d1634cc7-c2df-4015-bd8e-792616b6dc9b · outbound

This paper cites Towards Best Practices of Activation Patching in Language Models: Metrics and Methods.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:31:18.203317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:31:18.203317Z digest=sha256:076f5b9b97007b609acbe25248b01540b40e9dc0b8b46db3ccdfd6784d885ee0

Observation f9d3462e-56da-4ec7-8cf4-b8fc3231803d · outbound

This paper cites Transformers without normalization, 2025.

Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability Transformers without normalization, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:31:18.679971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-06T20:31:18.269637Z digest=sha256:d1fa1bf4b620c2618f4ad66d7ea67277b6b32e3a76ad42a650ce98ad0c2f2313

Pith citing papers

Observation c1b618d1-fb46-4671-98c7-5601df990771 · inbound

Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers cites this paper.

Key and Value Weights Are Probably All You Need: On the Necessity of the Query, Key, Value weight Triplet in Self-Attention Transformers Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:40:50.558061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T03:38:36.932424Z digest=sha256:abe135f345993dbe2814a7cfa3ec016fb4330d29662a7975f379a3c3fae630d3

Observation 84b8c5e9-ed7f-4d58-a553-8e19648e3cd2 · inbound

Discovering Interpretable Algorithms by Decompiling Transformers to RASP cites this paper.

Discovering Interpretable Algorithms by Decompiling Transformers to RASP Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T03:11:43.051506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:11:43.051506Z digest=sha256:b74757aa9f8f001c9daf3b46f3332608548d23440efb1236667d0a20de0feedf

Observation 23ef7d84-e166-4f62-8de0-541ef57c2998 · inbound

Gated Normalization Removal and Scale Anchoring in Pre-Norm Transformers cites this paper.

Gated Normalization Removal and Scale Anchoring in Pre-Norm Transformers Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:20:13.540889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T14:19:46.400753Z digest=sha256:3ae3669262ef09c7658a4a94ecd1df6cf98755f1bb2fc7d3bb817f2438b87e71

Observation 5ac88c08-06e6-441e-ba65-605d304cf889 · inbound

Selective Neuron Amplification in Transformer Language Models cites this paper.

Selective Neuron Amplification in Transformer Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:25:52.362679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:31:17.182481Z digest=sha256:d444b21b9f5a6a66b44599f413dc55bfdd3ebb942b84aacdd851e40d9b3bbd1b

Observation 42cc7a73-a64c-458f-9f99-2e469bff0c8f · inbound

Selective Neuron Amplification in Transformer Language Models cites this paper.

Selective Neuron Amplification in Transformer Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:16:25.133581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:28:05.507808Z digest=sha256:a99c2e8af5f236d2ecacb266f025c4eb59768eee161210cbc9cf987c7e5f08c3

Observation a0fa120d-8d23-4b81-b0b6-15533262c7e2 · inbound

Towards Verifiable Transformers: Solver-Checkable Circuit Explanations cites this paper.

Towards Verifiable Transformers: Solver-Checkable Circuit Explanations Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T15:05:48.064804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:47:13.178911Z digest=sha256:c72720349ce927fcacc132c4ce84d08ab4e3ff839a5494bb47297567643701b7

Observation b342e9cf-cb1f-4e01-bfaa-8a2cddd8a182 · inbound

Towards Verifiable Transformers: Solver-Checkable Circuit Explanations cites this paper.

Towards Verifiable Transformers: Solver-Checkable Circuit Explanations Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T13:29:53.932972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:29:53.932972Z digest=sha256:d00d733c38b9c691ef421f974136ecfc73c99840d0bb30b2150bffe28357ce54

Observation a357d3c0-3e46-4594-874e-a9d86d4d0a92 · inbound

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models cites this paper.

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T17:35:51.306825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T03:35:34.594617Z digest=sha256:817c0512dd1e424e695bab4d32f13589e7b30f557f4587d32cadfb8d5f5604ef

Observation 7dc6d4e6-4bc0-4830-a4e2-274b3eaefe57 · inbound

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models cites this paper.

Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models Transformers Don't Need LayerNorm at Inference Time: Scaling LayerNorm Removal to GPT-2 XL and the Implications for Mechanistic Interpretability

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T09:34:34.461654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T09:32:59.824110Z digest=sha256:d4ee0c7e8a5e5420c34747df1091d0c7437682045d0f55e2bbb60298949fe151