Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T21:26:42.281399Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2411.08790.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T21:26:42.281399Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:20:01.640542Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T15:20:08.711872Z
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation abb05878-58e8-4ea9-935f-d3ea7f231ea4 · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Foundational Challenges in Assuring Alignment and Safety of Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f81a1ebb-29a5-44bc-9521-7ac6facaa440 · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Refusal in Language Models Is Mediated by a Single Direction
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81177b00-329f-4445-923b-a851bd84995b · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Towards monosemanticity: Decomposing language models with dictionary learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71c6461e-4406-48b8-96ba-6adb713ca091 · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Progress update #1 from the GDM mech interp team
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e0219068-06dd-4abb-a94b-5075e75f8a87 · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Sparse Autoencoders Find Highly Interpretable Features in Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 033856ad-c918-4231-8177-75aaf694e5fd · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7acb2adf-8935-4b29-b4b0-ced4b2456b51 · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Scaling and evaluating sparse autoencoders
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99106a46-df2c-4d72-b85a-c52220f51816 · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Extract- ing sae task features for in-context learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2a9da4f6-2dee-4565-876d-b723b2fe74df · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec9e69e6-7462-4872-b81b-c07493e66d53 · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dcede98-43ca-41d7-84a4-ed6c5d787e5d · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 250e0322-6fb2-4e22-9062-897dfbbd4407 · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 858c7074-c469-4184-a78d-05176e644844 · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Steering Llama 2 via Contrastive Activation Addition
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbd93b3a-c248-4dd1-a11f-8dfe341ab134 · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Discovering Language Model Behaviors with Model-Written Evaluations
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aef3e9f5-2e52-4bf2-a019-fb1725854d4d · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f28daf5-ece9-4779-9f3b-0dcfc11a073e · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Progress update #1 from the gdm mech interp team
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cee9464a-f715-447f-8ee0-9968495d0161 · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Steering vectors github, 2024
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 926a9237-9662-46cd-8027-eb0cfc23c41c · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Analyzing the Generalization and Reliability of Steering Vectors
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d2a6ff2-a00b-4eb4-9098-40e0958cd4ff · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation aa9d9698-8add-4c34-8774-58ada3ff1f26 · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Vazquez, Ulisse Mini, and Monte MacDiarmid
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5d45b20a-0088-4159-adf1-1876ebabf847 · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfc2846d-b426-4898-b074-e6399d138de9 · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Steering Language Models With Activation Engineering
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e66b072-cd05-466f-a240-4a78e333966e · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Representation Engineering: A Top-Down Approach to AI Transparency
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5582750b-97be-46e4-916b-82418641a8ce · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Extending Activation Steering to Broad Skills and Multiple Behaviours
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e666d101-a9dd-4268-be4a-df728059110b · outbound
Can sparse autoencoders be used to decompose and interpret steering vectors? Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7efb8429-131f-4ba5-9ae0-ed4562b10300 · inbound
Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs Can sparse autoencoders be used to decompose and interpret steering vectors?
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.