Pith. sign in

Paper Citation Record · LEDGER

Can sparse autoencoders be used to decompose and interpret steering vectors?

As of 14 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2411.08790.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.08790 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:26:42.281399Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:20:01.640542Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:20:08.711872Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation abb05878-58e8-4ea9-935f-d3ea7f231ea4 · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Can sparse autoencoders be used to decompose and interpret steering vectors? Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.160078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.160078Z digest=sha256:7a6ade7d68da4ceb0456bc6bcbc7b4dcd068ccc49c0a35bffc99faaa7cdf2567

Observation f81a1ebb-29a5-44bc-9521-7ac6facaa440 · outbound

This paper cites Refusal in Language Models Is Mediated by a Single Direction.

Can sparse autoencoders be used to decompose and interpret steering vectors? Refusal in Language Models Is Mediated by a Single Direction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.165950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.165950Z digest=sha256:470d76bcf471c33fe0b36054a6df5b61fa9528a72b8ec09c5687302ff9e83b43

Observation 81177b00-329f-4445-923b-a851bd84995b · outbound

This paper cites Towards monosemanticity: Decomposing language models with dictionary learning.

Can sparse autoencoders be used to decompose and interpret steering vectors? Towards monosemanticity: Decomposing language models with dictionary learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.171114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.171114Z digest=sha256:2b0ac8375712a0161e2a5a511f4b65eecffb1fb630b2da695f2430c6cbdf1d7f

Observation 71c6461e-4406-48b8-96ba-6adb713ca091 · outbound

This paper cites Progress update #1 from the GDM mech interp team.

Can sparse autoencoders be used to decompose and interpret steering vectors? Progress update #1 from the GDM mech interp team

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:26:42.695410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:26:42.176621Z digest=sha256:df7a39c614e9af3ba9bd164972e579a66166e0b96f2dd1f297ca5663d4d77456

Observation e0219068-06dd-4abb-a94b-5075e75f8a87 · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Can sparse autoencoders be used to decompose and interpret steering vectors? Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.181551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.181551Z digest=sha256:c4677b714c9c53566ce6634ef20d558c8978582b72d6fd486423b5e65926b731

Observation 033856ad-c918-4231-8177-75aaf694e5fd · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Can sparse autoencoders be used to decompose and interpret steering vectors? The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.186908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.186908Z digest=sha256:fc02aafd9a3cb075095ee3f4651624d974870fea528cabc70679b31096a95c98

Observation 7acb2adf-8935-4b29-b4b0-ced4b2456b51 · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Can sparse autoencoders be used to decompose and interpret steering vectors? Scaling and evaluating sparse autoencoders

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.192559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.192559Z digest=sha256:34c813d7f11b8bdf8f3dd189980066ed0fd3df1fa59bfb54dd4ce51d7c44d7ef

Observation 99106a46-df2c-4d72-b85a-c52220f51816 · outbound

This paper cites Extract- ing sae task features for in-context learning.

Can sparse autoencoders be used to decompose and interpret steering vectors? Extract- ing sae task features for in-context learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:26:42.678272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:26:42.197485Z digest=sha256:ceecd9209088362d7099494f0ebf0fd9d9f9019701a4848aa7c85be3f7a9dfe3

Observation 2a9da4f6-2dee-4565-876d-b723b2fe74df · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

Can sparse autoencoders be used to decompose and interpret steering vectors? Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.207022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.207022Z digest=sha256:70dfa24a5e41eb445d42f6b3b5e1eab05c01e511f4f0bd850fe7d2896dee6b99

Observation ec9e69e6-7462-4872-b81b-c07493e66d53 · outbound

This paper cites In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering.

Can sparse autoencoders be used to decompose and interpret steering vectors? In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.212109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.212109Z digest=sha256:1d000d870ba3af86fe51ccecf995921c4e00268ed0d973ab736b52289f5903e3

Observation 5dcede98-43ca-41d7-84a4-ed6c5d787e5d · outbound

This paper cites Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models.

Can sparse autoencoders be used to decompose and interpret steering vectors? Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.218154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.218154Z digest=sha256:53f2cb2c3a8fe98697af75ca0ae71ae1e322694e52efc41b2ec1151befb53fa9

Observation 250e0322-6fb2-4e22-9062-897dfbbd4407 · outbound

This paper cites Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models.

Can sparse autoencoders be used to decompose and interpret steering vectors? Sparse Autoencoders Enable Scalable and Reliable Circuit Identification in Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.223217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.223217Z digest=sha256:6ecca964c069d8d888f43cc6e567b23ebbc71af84d5b8dc34c8c61ac48f95826

Observation 858c7074-c469-4184-a78d-05176e644844 · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Can sparse autoencoders be used to decompose and interpret steering vectors? Steering Llama 2 via Contrastive Activation Addition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.228251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.228251Z digest=sha256:9eb98ccb89f43a46b0b140ad57382706c9623e3a877cd0ecf8709b3fcb5b27b5

Observation cbd93b3a-c248-4dd1-a11f-8dfe341ab134 · outbound

This paper cites Discovering Language Model Behaviors with Model-Written Evaluations.

Can sparse autoencoders be used to decompose and interpret steering vectors? Discovering Language Model Behaviors with Model-Written Evaluations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.233435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.233435Z digest=sha256:5c0953e95bba91af2aa8755a66633914e01bce471b8a75b3c2fdb90af7d38f13

Observation aef3e9f5-2e52-4bf2-a019-fb1725854d4d · outbound

This paper cites Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders.

Can sparse autoencoders be used to decompose and interpret steering vectors? Jumping Ahead: Improving Reconstruction Fidelity with JumpReLU Sparse Autoencoders

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.238427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.238427Z digest=sha256:fd2282e0c35aa7ed0b28e782f01fe732bb225a3db6aa72493772b3168f16cd0b

Observation 4f28daf5-ece9-4779-9f3b-0dcfc11a073e · outbound

This paper cites Progress update #1 from the gdm mech interp team.

Can sparse autoencoders be used to decompose and interpret steering vectors? Progress update #1 from the gdm mech interp team

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:26:42.646166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:26:42.243162Z digest=sha256:4d0ee08748a87f35c64efbe554fbb3c4fd21f5193ea7061d4272f64ac1b582ce

Observation cee9464a-f715-447f-8ee0-9968495d0161 · outbound

This paper cites Steering vectors github, 2024.

Can sparse autoencoders be used to decompose and interpret steering vectors? Steering vectors github, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:26:42.630077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:26:42.247751Z digest=sha256:ac2dfa60103aea38ec675859aca5bbab865b1f46fb75138629f7bd1a8a5136b3

Observation 926a9237-9662-46cd-8027-eb0cfc23c41c · outbound

This paper cites Analyzing the Generalization and Reliability of Steering Vectors.

Can sparse autoencoders be used to decompose and interpret steering vectors? Analyzing the Generalization and Reliability of Steering Vectors

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.252157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.252157Z digest=sha256:9216bb5f3821c41fbdbc07d397668c4e1555355fea9fab41ee8071e5a83b35f5

Observation 1d2a6ff2-a00b-4eb4-9098-40e0958cd4ff · outbound

This paper cites Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet.

Can sparse autoencoders be used to decompose and interpret steering vectors? Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:26:42.614375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:26:42.257201Z digest=sha256:06e9ba4f924cddf544176fff30f46b52a3ac8f22f52634be60cb89d3f940fa30

Observation aa9d9698-8add-4c34-8774-58ada3ff1f26 · outbound

This paper cites Vazquez, Ulisse Mini, and Monte MacDiarmid.

Can sparse autoencoders be used to decompose and interpret steering vectors? Vazquez, Ulisse Mini, and Monte MacDiarmid

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:26:42.597872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:26:42.262056Z digest=sha256:046fc84bad12449c49f03b757cf645ef61f278bc8e19b421041cff1f99fbd799

Observation 5d45b20a-0088-4159-adf1-1876ebabf847 · outbound

This paper cites Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity.

Can sparse autoencoders be used to decompose and interpret steering vectors? Model Editing as a Robust and Denoised variant of DPO: A Case Study on Toxicity

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.271629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.271629Z digest=sha256:de982d3ceb781a0cd42c2bc0faa4a4a0a11fcf90d6f3ecf3220c8abe429d261e

Observation cfc2846d-b426-4898-b074-e6399d138de9 · outbound

This paper cites Steering Language Models With Activation Engineering.

Can sparse autoencoders be used to decompose and interpret steering vectors? Steering Language Models With Activation Engineering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.266596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.266596Z digest=sha256:ddb146f3875b4d9b53cc14e21666d87a6fc647190cab6c2add06bc9421588ce6

Observation 5e66b072-cd05-466f-a240-4a78e333966e · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Can sparse autoencoders be used to decompose and interpret steering vectors? Representation Engineering: A Top-Down Approach to AI Transparency

Reference 23

Resolution
malformed identifier
no resolver link, observed 2026-08-12T21:26:42.281399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.281399Z digest=sha256:659032fe44bbae1201ddaa50ffb1cec42ce592009739cab5eabe243ab0a310a1

Observation 5582750b-97be-46e4-916b-82418641a8ce · outbound

This paper cites Extending Activation Steering to Broad Skills and Multiple Behaviours.

Can sparse autoencoders be used to decompose and interpret steering vectors? Extending Activation Steering to Broad Skills and Multiple Behaviours

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T21:26:42.276584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:26:42.276584Z digest=sha256:d8efda086858d281e1a6dd84c262ab8fe5b5a61408d1f5802ff1080d6d9ba14f

Observation e666d101-a9dd-4268-be4a-df728059110b · outbound

This paper cites an unresolved cited work.

Can sparse autoencoders be used to decompose and interpret steering vectors? Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-12T21:26:42.662573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:26:42.202242Z digest=sha256:2528499b61a6b7a74146babc10f4cd347346ca98208452245bd495f9aaf00f95

Pith citing papers

Observation 7efb8429-131f-4ba5-9ae0-ed4562b10300 · inbound

Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs cites this paper.

Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs Can sparse autoencoders be used to decompose and interpret steering vectors?

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:20:08.798186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T15:20:01.640542Z digest=sha256:6b3eeb7e43e3a8507b75f9df7910ded83b9eec07d357f48f4835fa38b1a9a810