Pith. sign in

Paper Citation Record · LEDGER

Can sparse autoencoders be used to decompose and interpret steering vectors?

As of 14 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2411.08790.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.08790 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:26:42.281399Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:20:01.640542Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:20:08.711872Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation abb05878-58e8-4ea9-935f-d3ea7f231ea4 · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Can sparse autoencoders be used to decompose and interpret steering vectors? Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.160078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.160078Z digest=sha256:da5c0436d8b80a633ceb840c3dd916a96eee08807e5042f93df0f158a7d5e62f

Observation f81a1ebb-29a5-44bc-9521-7ac6facaa440 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Can sparse autoencoders be used to decompose and interpret steering vectors? Refusal in Language Models Is Mediated by a Single Direction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.165950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.165950Z digest=sha256:ff7c852bc410e0ae4d49bc243c721336e7ee6653368cf85195da87e2eac70a8d

Observation 81177b00-329f-4445-923b-a851bd84995b · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning.

Can sparse autoencoders be used to decompose and interpret steering vectors? Towards monosemanticity: Decomposing language models with dictionary learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.171114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.171114Z digest=sha256:8cb335df5c1c79783e38b5eb23a9966402c1fb10712d992941cbb236d8855195

Observation 71c6461e-4406-48b8-96ba-6adb713ca091 · outbound

This paper cites Progress update #1 from the GDM mech interp team.

Can sparse autoencoders be used to decompose and interpret steering vectors? Progress update #1 from the GDM mech interp team

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:26:42.695410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:26:42.176621Z digest=sha256:c7d60a1b4bdacc8cc95ae8ba86b7c1f5db9b57eb240bb164d9cfd0496945eeda

Observation e0219068-06dd-4abb-a94b-5075e75f8a87 · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Can sparse autoencoders be used to decompose and interpret steering vectors? Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.181551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.181551Z digest=sha256:3b85e6171d92a548ec1331ebcd64e5a3f9247e182ed08627f1e552e2a9a771ae

Observation 033856ad-c918-4231-8177-75aaf694e5fd · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Can sparse autoencoders be used to decompose and interpret steering vectors? The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.186908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.186908Z digest=sha256:777ed404ec70e95529e2ceffd36d27f04bdc5a43fb33529d40aff7a1358113fe

Observation 7acb2adf-8935-4b29-b4b0-ced4b2456b51 · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Can sparse autoencoders be used to decompose and interpret steering vectors? Scaling and evaluating sparse autoencoders

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.192559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.192559Z digest=sha256:823248ded940f3e404276c9d8219940834a38cb384cdb777c5cde339b5a1ab27

Observation 99106a46-df2c-4d72-b85a-c52220f51816 · outbound

This paper cites Extract- ing sae task features for in-context learning.

Can sparse autoencoders be used to decompose and interpret steering vectors? Extract- ing sae task features for in-context learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:26:42.678272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:26:42.197485Z digest=sha256:848886cbe38cd895d77a4b5ba7350571846220ceb5e2cf829cf26e7c6002ad28

Observation 2a9da4f6-2dee-4565-876d-b723b2fe74df · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

Can sparse autoencoders be used to decompose and interpret steering vectors? Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.207022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.207022Z digest=sha256:c7d2b03d90e099c163ebfb527756a252ea879d4ff1362c2e18cf06d9ff65b22a

Observation ec9e69e6-7462-4872-b81b-c07493e66d53 · outbound

This paper cites In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering.

Can sparse autoencoders be used to decompose and interpret steering vectors? In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.212109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.212109Z digest=sha256:03aea9d3fe52c8b7fa8b679e859b8f436c2eda7c1162194d5d6f2001c27bc1c2

Observation 5dcede98-43ca-41d7-84a4-ed6c5d787e5d · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Can sparse autoencoders be used to decompose and interpret steering vectors? Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.218154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.218154Z digest=sha256:7f4c072938e486e4e47d8260de5d2e8e749afce07f59815e3e69767bc8b12100

Observation 250e0322-6fb2-4e22-9062-897dfbbd4407 · outbound

This paper cites Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models.

Can sparse autoencoders be used to decompose and interpret steering vectors? Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.223217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.223217Z digest=sha256:b6c10251d054169efd5a54ee999e8305d954b8e939e5013379911a68d36ae2cb

Observation 858c7074-c469-4184-a78d-05176e644844 · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Can sparse autoencoders be used to decompose and interpret steering vectors? Steering Llama 2 via Contrastive Activation Addition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.228251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.228251Z digest=sha256:51b3fa0c2972bc16928ec06df608f66cd69ace7f70a8b827fda7653ff45ae265

Observation cbd93b3a-c248-4dd1-a11f-8dfe341ab134 · outbound

This paper cites Discovering Language Model Behaviors with Model-Written Evaluations.

Can sparse autoencoders be used to decompose and interpret steering vectors? Discovering Language Model Behaviors with Model-Written Evaluations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.233435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.233435Z digest=sha256:421ca670603f346311390fc9e23d6cf350ed1a16cbf8b35f161b02afc4fd8d42

Observation aef3e9f5-2e52-4bf2-a019-fb1725854d4d · outbound

This paper cites Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders.

Can sparse autoencoders be used to decompose and interpret steering vectors? Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.238427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.238427Z digest=sha256:e4f2f9fbd67d1fe419473f38a5d0be53ee993c46384ca97f2fb00992c140c580

Observation 4f28daf5-ece9-4779-9f3b-0dcfc11a073e · outbound

This paper cites Progress update #1 from the gdm mech interp team.

Can sparse autoencoders be used to decompose and interpret steering vectors? Progress update #1 from the gdm mech interp team

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:26:42.646166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:26:42.243162Z digest=sha256:3a9abfcdb14e94ed9d71765ddab893ad9b4f08ac84d1e6767dc47c192e492ec2

Observation cee9464a-f715-447f-8ee0-9968495d0161 · outbound

This paper cites Steering vectors github, 2024.

Can sparse autoencoders be used to decompose and interpret steering vectors? Steering vectors github, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:26:42.630077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:26:42.247751Z digest=sha256:599e9287732ccec4eecb4ef8696b2299b62d9e170d18d16203d0fa32ba4a3046

Observation 926a9237-9662-46cd-8027-eb0cfc23c41c · outbound

This paper cites Analyzing the Generalization and Reliability of Steering Vectors.

Can sparse autoencoders be used to decompose and interpret steering vectors? Analyzing the Generalization and Reliability of Steering Vectors

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.252157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.252157Z digest=sha256:19cb58124903191efe084f1acfe1f9dc7778da54b1c1a419f4be4dc948edb68b

Observation 1d2a6ff2-a00b-4eb4-9098-40e0958cd4ff · outbound

This paper cites Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet.

Can sparse autoencoders be used to decompose and interpret steering vectors? Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:26:42.614375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:26:42.257201Z digest=sha256:ed4e3e1feb9c0431d4cccacd53df47147746db2a7619ec085e67f77ca90ba35f

Observation aa9d9698-8add-4c34-8774-58ada3ff1f26 · outbound

This paper cites Vazquez, Ulisse Mini, and Monte MacDiarmid.

Can sparse autoencoders be used to decompose and interpret steering vectors? Vazquez, Ulisse Mini, and Monte MacDiarmid

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:26:42.597872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:26:42.262056Z digest=sha256:2e0755b8802a6f3b8f103e8429480f9c204977eb4904660bfba492575e5aa496

Observation 5d45b20a-0088-4159-adf1-1876ebabf847 · outbound

This paper cites Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity.

Can sparse autoencoders be used to decompose and interpret steering vectors? Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.271629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.271629Z digest=sha256:1dd476f4f09cde92dddb16f63021e66f439db827632e8d6dce47da5edfd4f632

Observation cfc2846d-b426-4898-b074-e6399d138de9 · outbound

This paper cites Steering Language Models With Activation Engineering.

Can sparse autoencoders be used to decompose and interpret steering vectors? Steering Language Models With Activation Engineering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.266596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.266596Z digest=sha256:78451eb6402a781e86f3fd9e006ba50341ccaf37e4692d62f67ddc53c868d803

Observation 5e66b072-cd05-466f-a240-4a78e333966e · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Can sparse autoencoders be used to decompose and interpret steering vectors? Representation Engineering: A Top-Down Approach to AI Transparency

Reference 23

Resolution
malformed identifier
no resolver link, observed 2026-08-12T21:26:42.281399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.281399Z digest=sha256:9e8990fa81734a589e701480c1adba01c8aecabbd5169b5aa9788246636054d5

Observation 5582750b-97be-46e4-916b-82418641a8ce · outbound

This paper cites Extending Activation Steering to Broad Skills and Multiple Behaviours.

Can sparse autoencoders be used to decompose and interpret steering vectors? Extending Activation Steering to Broad Skills and Multiple Behaviours

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.276584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.276584Z digest=sha256:dbe52a79e7d8330846c721cad8977be3726259a263b36140f58bbaa954893fb2

Observation e666d101-a9dd-4268-be4a-df728059110b · outbound

This paper cites an unresolved cited work.

Can sparse autoencoders be used to decompose and interpret steering vectors? Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:26:42.662573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T21:26:42.202242Z digest=sha256:4258d44f7241a8e1a27c5375b9284f797aeacab431451fe468b5f97b6aff17a7

Pith citing papers

Observation 7efb8429-131f-4ba5-9ae0-ed4562b10300 · inbound

Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs cites this paper.

Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs Can sparse autoencoders be used to decompose and interpret steering vectors?

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:20:08.798186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T15:20:01.640542Z digest=sha256:feea8b4add342770f3bfbc7ff064261575b2155d6308fef17c80faed11ecb470