Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:28:16.065859Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 2 inbound Pith citation observations for arXiv:2501.06254.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:28:16.065859Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T18:47:45.375094Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T21:15:04.055124Z
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cc4e984b-33c7-4f71-9478-27ff0f584fda · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Mechanistic interpretability for AI safety - a review
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3b814fd9-f599-4004-823e-04e7403f66f0 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Pythia: A suite for analyzing large language models across training and scaling
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a08aa4bf-cca9-408f-a4a9-43899f3da6a1 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Towards monosemanticity: Decomposing language models with dictionary learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff5e80dd-be4e-44cc-9920-963f851b0309 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Evaluating Open-Source Sparse Autoencoders on Disentangling Factual Knowledge in GPT-2 Small
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b267948a-e41a-4928-89c5-7209fa478ef8 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Sparse Autoencoders Find Highly Interpretable Features in Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16a6db1e-75a5-46a5-8d12-d7e5299d1de1 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Interpreting and steering features in images, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4e7bb341-b664-412c-a0a2-7fd4bcdc3c0c · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Toy models of superposition
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 970fa9b9-db1e-4fac-a602-19671ceb57a0 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words JumpReLU: A Retrofit Defense Strategy for Adversarial Attacks
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02ed482f-ad9e-42c3-aded-4d4ca215d020 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Towards Empirical Interpretation of Internal Circuits and Properties in Grokked Transformers on Modular Polynomials
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81e8df3a-b4bc-4ca7-b374-d7c5599d98db · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Scaling and evaluating sparse autoencoders
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0bcaa25-0bea-4321-b351-0b0694d112d8 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Gemma: Open Models Based on Gemini Research and Technology
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f2cd971-2d58-43c8-9e55-cc1d6327c0e9 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Finding Neurons in a Haystack: Case Studies with Sparse Probing
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 297934e4-6121-4421-9d77-4589a307d0f2 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edec0ad7-d7bf-44d3-a7ab-012dca57712a · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Ghost grads: An improvement on resampling
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4711a0d7-6c4a-4d35-a5bd-14bba94ada80 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Saebench: A comprehensive benchmark for sparse autoencoders, 2024 a
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0debf41d-2271-46f6-884f-9bef3784adb7 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Evaluating Sparse Autoencoders on Targeted Concept Erasure Tasks
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7610f609-cf6d-4d13-aada-894e10abdfa1 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Interpreting Attention Layer Outputs with Sparse Autoencoders
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83307935-0e9c-43dd-a3b5-523394c7b517 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Gazing in the latent space with sparse autoencoders
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e84e1afb-5a51-423f-962d-6d65c12a9ae2 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words The Geometry of Concepts: Sparse Autoencoder Feature Structure
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b30ba452-01ef-4467-ba58-f27932b2d58d · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de2a5c91-fb68-4c4e-b887-448a37b15947 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f15ea0c9-eef9-4353-b38c-3cdb31ab3af3 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Towards Principled Evaluations of Sparse Autoencoders for Interpretability and Control
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a79cda67-dc31-4687-ba63-98afb01110d1 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Bridging Lottery Ticket and Grokking: Understanding Grokking from Inner Structure of Networks
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e7dc628-b26d-43a5-bd8a-16342189fa84 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Progress measures for grokking via mechanistic interpretability
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13744a5e-33de-4ac7-85e9-5c9ad6bb8fa8 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words interpreting gpt: the logit lens
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1cef830c-a104-4898-8b9c-56cf0b1aa796 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Zoom in: An introduction to circuits
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation abe6630a-74d8-4096-9ae8-a8032ec724da · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words WiC: the Word-in-Context Dataset for Evaluating Context-Sensitive Meaning Representations
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37181764-fde1-408a-a9b1-0f989ed3b46f · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9cb7fd5-a01b-4878-9af2-bb33f1531de9 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Language models are unsupervised multitask learners, 2019
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5bf6374-ad56-4d9c-aa5f-43fb62817d81 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words XL - W i C : A multilingual benchmark for evaluating semantic contextualization
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 745d2a4b-74ce-4277-bec0-97532b2831b9 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Improving Dictionary Learning with Gated Sparse Autoencoders
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 596a6086-14ec-47f8-800c-39438f1e92d4 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9768eeb9-2c9a-433d-8cc7-f2c85ae6c26d · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Interpreting preference models w/ sparse autoencoders, 2024
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f7f4036e-b3a8-4fc2-bbc5-3f5eeb9a67dd · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words IllusionVQA: A Challenging Optical Illusion Dataset for Vision Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41906680-b897-405f-b028-15bd6191c3e0 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Taking features out of superposition with sparse autoencoders, 2022
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5d953090-0072-46ce-8402-4facfbf74a64 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words On the proper treatment of connectionism
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 06033563-db10-4f6b-899c-b5c470fe7559 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Unpacking sdxl turbo: Interpreting text-to-image models with sparse autoencoders
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2956a9d6-c3cc-4d02-88a6-69e2ddd75bf9 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Daniel Freeman, Theodore R
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca6413c9-4eec-4a01-821a-2a6ab10e8abd · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 506f5d05-097b-47e5-8de2-c7c498d0d4b7 · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64baeab7-5868-40c3-b5ed-404054ea7e1a · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words RedPajama: an Open Dataset for Training Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 406ead5a-3daa-4ab3-96b1-6ae4ebe401dd · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words write newline
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad7fbe71-11e6-483d-b09a-89ff751fc05d · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words @esa (Ref
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c407f0f-5cda-43fc-b145-314012176faf · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 517a21a2-09b3-465a-9577-f83bf1f7a06f · outbound
Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff612882-61eb-4bd9-a2c0-f3583fd53bc3 · inbound
Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
Reference 105
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77fa3917-5a51-404d-a2af-bb338a0d669a · inbound
The Rate-Distortion-Polysemanticity Tradeoff in SAEs Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous Words
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.