Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T21:12:13.917656Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 23 inbound Pith citation observations for arXiv:2502.04878.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T21:12:13.917656Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:03:00.843301Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
30 of 30 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 874c877f-280d-497d-bf95-e3984e979e07 · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Transcoders Find Interpretable LLM Feature Circuits
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 239891e1-7ec4-42f5-adef-30863cf8261e · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Scaling and evaluating sparse autoencoders
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d22a8e1-406f-4e90-b0c1-ff1f72e12c7a · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Dissecting Recall of Factual Associations in Auto-Regressive Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42ee1fd8-7de2-4baf-b442-6783740ff9ee · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Language Models Represent Space and Time
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f60d5c5-bb85-431f-a961-bef702da7580 · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Interpreting Attention Layer Outputs with Sparse Autoencoders
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3815e31b-4fd3-4f2d-84e6-28efbfb6b15c · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc6ddbfc-4f97-4441-a16a-813e4a5fcc4a · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9ba40f5-e078-4ae5-a2de-b448ab61c189 · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Locating and editing factual associations in gpt
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2e5bcfbd-e44b-40ce-8d4d-92ad719e56e0 · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Progress measures for grokking via mechanistic interpretability
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49d29ef5-c3b9-4962-a2b7-4dcc3cb0f65e · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis However, in practice we find that the bdecs of SAEs trained on the same latents are very similar (minimum cosine similarity of 0.9970, differing by less than 0.1% in magnitude)
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ef503a3c-c93f-4473-9a87-b88950e81bfc · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Gemma 2: Improving Open Language Models at a Practical Size
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ab72c83-bf32-4747-89cd-887135d947a0 · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e21bddf-f330-44fe-8cb8-15dc0fc2d78f · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Relational Composition in Neural Networks: A Survey and Call to Action
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2703105c-501b-43ed-8254-cd148002a0e0 · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Typically measured as average L0 across a batch: L0 = 1 n P i ∥f (xi)∥0
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6b01bdb2-aa42-4233-967c-8e11c63c0211 · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis This provides a gradient for training unlike the L0-norm, but suppresses latent activations harming reconstruction performance (Rajamanoharan et al., 2024a)
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 75b836a2-3e6f-43ba-81ea-97bafebf99b6 · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis make sure
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5b66ee5d-585e-48f4-ad4c-3bbdc2b858b6 · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis A.5 O PEN SOURCE SAE WEIGHTS All GPT-2 Small SAEs were trained on the layer 8 residual stream, which was chosen in line with Gao et al
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 192fd351-476e-47c5-9b4c-101350f33c37 · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis We used the TransformerLens ( https://transformerlensorg.github.io/ TransformerLens/) implementations of GPT-2 and Gemma 2 2B
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4033a41d-75f9-4879-bd62-79ba8f998c96 · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 730dedff-5172-4142-a060-29ca0ecbe48a · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Figure 21: Sparse probing evaluation accuracy by GPT-2 SAE dictionary size across 8 benchmark datasets, with a sparse probe using the top latent
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9bf66e8c-47ce-4693-8b75-852b380bbea2 · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis 22 Published as a conference paper at ICLR 2025 Figure 23: TPP evaluation accuracy by GPT-2 SAE dictionary size across 2 benchmark datasets, ablating up to 50 latents
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 77c682d5-d3a9-459b-8f49-dee665636ed8 · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis We find a lower threshold for distinguishing novel features from reconstruction features (0.4)
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 78cdf7f9-3d16-4384-ac2e-7458f9175793 · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis In-context Learning and Induction Heads
Reference 1997
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11610842-c7ca-4265-99e0-90cde01e776d · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis k-Sparse Autoencoders
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5f1cc76-a6a5-46fd-b5e3-13ef2965bb31 · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 409d70f5-73cc-4304-8ba0-9bde432a44d0 · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Improving Dictionary Learning with Gated Sparse Autoencoders
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05650c7c-ca77-4ece-abde-b43b5a1f9dd9 · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Toy Models of Superposition
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53139257-4cc2-4178-a89a-e7ad60874756 · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Not All Language Model Features Are One-Dimensionally Linear
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fe28006-0be0-4138-b8f7-57db0079805f · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis BatchTopK Sparse Autoencoders
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0916881-ecaa-4c01-bb3a-c50a8aa1698f · outbound
Sparse Autoencoders Do Not Find Canonical Units of Analysis Sparse Autoencoders Find Highly Interpretable Features in Language Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9472c785-10ea-4cc4-bee1-c5ab6b941e6d · inbound
Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50f4f490-74c5-4fce-b7d5-385ecd300fd4 · inbound
Stochastic Parameter Decomposition Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 158f39e2-2786-4adc-92dc-0e2a90c0ce6f · inbound
Teach Old SAEs New Domain Tricks with Boosting Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 899a1934-92c6-408c-95e4-b7e8a9da13b8 · inbound
Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f241173c-da4b-4c12-9b56-37ad8c59b6ac · inbound
Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb0f0e78-985c-4cfb-9741-7beb6e1dfbe4 · inbound
Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 383309bf-692f-4e12-bf36-4bc3435bf553 · inbound
Superposition Without Interference? Towards Isolated Interventions via Almost Orthogonal Features in Language Models Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7048f3de-b19f-4ae9-af2e-5389e9cd3515 · inbound
Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7f31fb58-494e-45ad-bc04-5f179e4536d7 · inbound
Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dab9fcd5-3d2e-4065-85c4-0899531f3624 · inbound
Position: Mechanistic Interpretability Must Disclose Identification Assumptions for Causal Claims Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1fecbe30-0288-4573-b4b9-ffb45395717b · inbound
Disentangled Sparse Representations for Concept-Separated Diffusion Unlearning Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 636848fb-09da-4f38-bd53-68091360e0b3 · inbound
WriteSAE: Sparse Autoencoders for Recurrent State Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ccd7079f-fd64-42cb-adf8-1bdc9e4c1493 · inbound
WriteSAE: Sparse Autoencoders for Recurrent State Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 601d00b0-efa3-4301-95ad-071b3edf7a0b · inbound
WriteSAE: Sparse Autoencoders for Recurrent State Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0a3b2eb0-d68d-43ae-a24c-623bd614bbb4 · inbound
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 11339158-52c7-4c72-86ab-7a76859a9cf6 · inbound
Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cd88b479-ee21-454e-8851-c3ec323e2127 · inbound
Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69602583-7167-40d1-924a-e5a348149c10 · inbound
Size Doesn't Matter: Cosine-Scored Sparse Autoencoders Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation efad8060-df6f-4968-a4f5-5f023a01cae4 · inbound
Critical Percolation as a Synthetic Data Model for Interpretability Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cd273a25-47d3-499a-9d18-b24ccc51c51e · inbound
Do Sparse Autoencoders Learn Meaningful Concept Hierarchies? Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d75a73d9-4b33-479c-9ca4-50e195db219f · inbound
At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 51ddf6ad-60a4-4f2e-8637-86b23e9d3cc2 · inbound
Surrogate Fidelity: When Can Open LLMs Explain Closed Ones? Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 322759d2-a828-403a-89b9-196fa3ae6ba1 · inbound
Verbalizable Representations Form a Global Workspace in Language Models Sparse Autoencoders Do Not Find Canonical Units of Analysis
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.