Pith. sign in

Paper Citation Record · LEDGER

AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 45 inbound Pith citation observations for arXiv:2501.17148.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.17148 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 45 of 45 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:29:34.141936Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c213a571-1f47-4cf7-bf24-0257b25e7655 · inbound

Toward universal steering and monitoring of AI models cites this paper.

Toward universal steering and monitoring of AI models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T01:06:10.147605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T01:06:10.147605Z digest=sha256:9101ffea6046357c1d56df6776d7843f2be5acda3925f9d1b5887e57e329ec50

Observation 11c5b09b-adb1-4bb7-8871-25e75ff5c78f · inbound

Disentangling Polysemantic Channels in Convolutional Neural Networks cites this paper.

Disentangling Polysemantic Channels in Convolutional Neural Networks AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T12:29:34.141936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:29:34.141936Z digest=sha256:a5662095f15b0959e2f35c427b74b8638bff7f2a335c818e0704dd07295c8392

Observation e4670055-fc91-487f-9650-02c24615c8a9 · inbound

MIB: A Mechanistic Interpretability Benchmark cites this paper.

MIB: A Mechanistic Interpretability Benchmark AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:44.465708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:19:44.465708Z digest=sha256:6b846fdd0431ba71f4550ced4f8c5fa053d4756a13a3e2e356f249b6d2184a37

Observation 543ff38d-cc52-4e58-ae80-9fa7891252f6 · inbound

EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models cites this paper.

EasyEdit2: An Easy-to-use Steering Framework for Editing Large Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T11:35:04.755554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:35:04.755554Z digest=sha256:2e3b7502b98393976f6f54f93f030f7198f03315cbc576086403a26daa33d0ab

Observation 74a24061-0e17-4599-a309-ec652df7de8d · inbound

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii cites this paper.

Evaluating Explanations: An Explanatory Virtues Framework for Mechanistic Interpretability -- The Strange Science Part I.ii AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-16T04:25:06.093358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:25:06.093358Z digest=sha256:f779efc66daaa5a4de8b677bb985d6b8180f12a037716dbdf587e9a6ad1f30aa

Observation 7a9fc4e6-95b4-4fb6-a4ba-502bb7e25f20 · inbound

Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering cites this paper.

Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:30:22.282961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:30:22.282961Z digest=sha256:a3d09f4580823b8ef89709978fb550189924334f62e61f85285dd3efb06c4463

Observation 31cea893-d830-4649-a4bb-38ba2a5f3b78 · inbound

Steering Large Language Models for Machine Translation Personalization cites this paper.

Steering Large Language Models for Machine Translation Personalization AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:51.732241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:01:51.732241Z digest=sha256:afe0d1569b6a69cc31a5ecc815b0dcb59f8c8fe4cc44054bc953a105cd4bf53c

Observation f67b55c0-fde2-4415-afdd-9906c7cd1c92 · inbound

Evaluating Steering Techniques using Human Similarity Judgments cites this paper.

Evaluating Steering Techniques using Human Similarity Judgments AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:58.857369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:58.857369Z digest=sha256:f985b410761a50459ff21bc308788bf3d4e2f1ef1d6995c7a41fc22e504002eb

Observation 0bbd5b28-16ef-422c-85fd-6ec1fde77f96 · inbound

Improved Representation Steering for Language Models cites this paper.

Improved Representation Steering for Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:33.503653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:33.503653Z digest=sha256:525569848e77c8a3e3ee853325d910176d63ef7b4447918c57e83b92ede4f807

Observation 23e4f887-4e7a-45ad-ad63-9c1688b592c9 · inbound

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race cites this paper.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.564287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.564287Z digest=sha256:0882ad576af2bce3c4ed25d72fc0a5900b97118becfd811db52068ca4f656f63

Observation a44a648d-3daf-4ca6-80dc-25d7bcf6f345 · inbound

Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models cites this paper.

Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:11.939328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:07:11.939328Z digest=sha256:31df9d56148cc427399ae02ceb5c02dd61f701610a94668516f1e536a9f5d2a2

Observation de4cbd8d-1c7a-408d-9bbd-93ecc67f2526 · inbound

HyperSteer: Activation Steering at Scale with Hypernetworks cites this paper.

HyperSteer: Activation Steering at Scale with Hypernetworks AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:19.096812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:11:19.096812Z digest=sha256:3ba88fee35180eb0b28f963a5139d87bf1779289298146db3a3e1a195ed5ffc3

Observation 8f410453-9292-436c-a201-524a7700f897 · inbound

Fine-Grained Interpretation of Political Opinions in Large Language Models cites this paper.

Fine-Grained Interpretation of Political Opinions in Large Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:16.702496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:16.702496Z digest=sha256:e1ae9b984cb009009b721ba0896e7edb79bd29b692a0a1242298be31c80c6626

Observation 48448f35-3dc6-4cb3-b0e2-fd117fefe50a · inbound

Resa: Transparent Reasoning Models via SAEs cites this paper.

Resa: Transparent Reasoning Models via SAEs AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:42.741450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:42.741450Z digest=sha256:ab5ad73ffe3d943db6fc51ddbdcc7d99d07f99cbecbb4b4a128247475e93a1cc

Observation aa3cfd68-4461-42d7-9c7b-067819b1df5e · inbound

From Emergence to Control: Probing and Modulating Self-Reflection in Language Models cites this paper.

From Emergence to Control: Probing and Modulating Self-Reflection in Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T01:03:41.633267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:03:41.633267Z digest=sha256:48b21e5c202172bc0f5a4d6d8a787e0d04717fc2df80c3adb421790f4ccecd73

Observation 686e0ed0-7129-44b5-add7-9373dae8b45e · inbound

Position: Use Sparse Autoencoders to Discover Unknowns cites this paper.

Position: Use Sparse Autoencoders to Discover Unknowns AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:34:33.631283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:34:33.631283Z digest=sha256:b83c11659effbfaadf3deb34bd2b5998d926c4eb4ce1ff581cd668d99ec8bab8

Observation ecec1d06-3b49-41fc-859c-0a010f994c08 · inbound

Insights into a radiology-specialised multimodal large language model with sparse autoencoders cites this paper.

Insights into a radiology-specialised multimodal large language model with sparse autoencoders AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T16:40:44.637794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:40:44.637794Z digest=sha256:30eeafcdea27b812098e05ac30bc9277760e3779bd13523ded7df967560e7c5a

Observation 5f90aaef-7103-4c7f-a616-dc6de6694890 · inbound

Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders cites this paper.

Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T17:18:55.749890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:18:55.749890Z digest=sha256:840c8b8d7e6689e19354c068aab97ab0b8e2f052b4929a8b91ba8ffeb42cd42b

Observation 3d6046e9-9891-47b7-b248-19a4ac46b9a4 · inbound

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation cites this paper.

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-05T18:47:45.341219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:47:45.341219Z digest=sha256:91312847fb0ad4a1ed6c36820119049b92fbb28901a59ae773bebd85fe285469

Observation 64bae72e-6f2d-45e3-a86e-bcd48d11aca8 · inbound

Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA cites this paper.

Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T22:18:44.791907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:18:44.791907Z digest=sha256:e2afdc6299891ac6578be54f3f043cc2f5649dc575164fa835ede00dbbf50c3a

Observation fef9a84f-9bd2-402d-979e-78f556efb848 · inbound

VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models cites this paper.

VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:27.597133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:27.597133Z digest=sha256:7a33a9fe5c01b4d2a7546df3c27153b907de9341250c4393270a407f9ad2ce3d

Observation 6a29067e-97b4-4317-a117-068745e3b2cc · inbound

The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail cites this paper.

The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T18:33:09.003812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:33:09.003812Z digest=sha256:c50a20007237cd0418895dc0938b158e68e6186346ce255a72ba10fb59f0636c

Observation 25cc1439-9855-45b6-90e2-52c1d205dc7e · inbound

Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes cites this paper.

Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:31:10.587655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-08T11:42:16.090162Z digest=sha256:207fbf824aabc5cc5de9a5be2b7c0625e39c1fda3d37ab4defe6827f01a349ae

Observation 999f74ea-7573-496a-b670-863bce46d67f · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:51:15.332661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T00:50:06.234399Z digest=sha256:eff3c3af31880b7fffb68ec61b9f718b19fec41c03f3686fbfad6f32d9a1649a

Observation c9542428-9d16-464c-9340-0a67098a26b5 · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:22:58.969437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T21:21:37.185232Z digest=sha256:07c71606e08485440f2b0c46674622eed6bb8bd3621c312264b2fb4ccfdb9034

Observation cbcdd5f3-9ca9-4d3e-921f-0895e37a543a · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:07:41.693739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T17:06:03.483247Z digest=sha256:6c71552251e74e72b8e03cb3d3de5ba27add184e0c2aedc4910a9b96424e7585

Observation dc19f809-2c50-4736-a7d6-263c18fb155f · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:45:08.055382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T23:41:36.199099Z digest=sha256:f98e3700fa655c82c5a3086ce4c1f9688e6ee3feade9370d618997f69afe26ea

Observation cf6380e4-61c1-493d-9226-07ebcd27f71b · inbound

The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime cites this paper.

The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:11:26.524548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T03:37:12.567351Z digest=sha256:f2393bee026d7451b937474c43d5047dcf85b60c537812d7e271be71ef77179c

Observation 03000c2e-bb75-415a-8a01-5338990b0c2b · inbound

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space cites this paper.

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 109

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:22:18.867569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-13T05:17:34.283917Z digest=sha256:7b6ec73443478597acf17204ae5dc8db2213b57e0555dc58bdd4a5b3ceb5b7d8

Observation 5273eeed-d53c-4b66-aa4e-adcbe3a4aacb · inbound

WriteSAE: Sparse Autoencoders for Recurrent State cites this paper.

WriteSAE: Sparse Autoencoders for Recurrent State AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:49:50.175124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T07:46:41.159688Z digest=sha256:47c30d46a9a4a5214b50b031d285da7d10a3b65ee3f6a166ba6da06c67cb19cb

Observation d5d1b455-8957-4939-afef-686d00f5552e · inbound

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search cites this paper.

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:39:10.652456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T22:36:12.466550Z digest=sha256:a4eebbc404029b4a852c74f67f11469c76482b9c066a9a61af9e27f542837b9a

Observation 420f821d-c186-40f2-8e77-c80e1bd8075a · inbound

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search cites this paper.

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.537374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T09:59:32.201925Z digest=sha256:dc11a431fadafbbe6fb751e158335695510d16e7f512bb1ab53f7baa84389ab9

Observation 037a966f-e60b-465f-aa09-3c95ed64ef6f · inbound

Are Sparse Autoencoder Benchmarks Reliable? cites this paper.

Are Sparse Autoencoder Benchmarks Reliable? AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:43:16.805970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T12:43:13.014365Z digest=sha256:04118f80e6a87bc5f3982914bfa08c663a8b369fbc4bd1bddc88127c25ec4e7d

Observation f63048b1-1502-4a22-91ff-0f9842ba83ec · inbound

Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection cites this paper.

Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:03:29.907469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T13:53:27.306664Z digest=sha256:38060c525f9522e1664a155556463ed5bcb61ebb7062768a385332306cc2a7dc

Observation 8e55ea6d-c6f2-4dd1-a1a5-9344dddd060a · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:47.199446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T22:16:47.743387Z digest=sha256:ce0cc1762d2431c93c8b6acccd903c670cb1281b2c543bb28eb420f3f2b33149

Observation 1b40c27b-9c83-451e-bbbd-1341a9ee321f · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 117

Resolution
unresolved
no resolver link, observed 2026-07-12T17:03:44.315006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:03:44.315006Z digest=sha256:cafc6db4f28ca587a36ba3457502e609344afe1d7f8725a45a74272bca15af76

Observation e3ba96c8-ca5b-4522-8973-6c0e25c6d87e · inbound

When is Your LLM Steerable? cites this paper.

When is Your LLM Steerable? AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:27:56.578484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T10:01:21.682285Z digest=sha256:b2068f13c75a138aa792d45b5d30858ae8bd3c897ffc8b9cda97e52997f5e3b9

Observation e79bbe08-8e29-4e61-8e50-6f9acde5ce5b · inbound

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal cites this paper.

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 105

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:07:47.864621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T10:32:57.295159Z digest=sha256:f15e9b3003997a42262960d189086b14b70f43358f699db968a325072db22d3c

Observation 71cc1d01-e169-4ee8-bfb5-708a7379f935 · inbound

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders cites this paper.

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:15:29.624121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-01T07:15:16.674714Z digest=sha256:323ff8bf5cc8878b88312566ed794dae35c9837431dbe7480564d58e9de78ef4

Observation 67b37175-3b56-4910-984d-8f416ca019f6 · inbound

Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent cites this paper.

Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T20:58:54.009938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:58:54.009938Z digest=sha256:8bac73694a47b639e48b54c1d8de209997a4bb342ceadc5d29d841e294c6d7f2

Observation 608fb2c6-208f-4036-b5af-45253a520b6a · inbound

Probabilistic Concept-Aware Steering for Trustworthy LLM Inference cites this paper.

Probabilistic Concept-Aware Steering for Trustworthy LLM Inference AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-02T13:56:50.382262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:56:50.382262Z digest=sha256:0ed93270986f5a586b763c64efc70386e73adaef37d7b9c5a42d2c2e6d55a77d

Observation 7ee8c3dc-41ae-45d0-9aa5-9d79a82c5a87 · inbound

Inverted Detection and Control in Steering Vectors cites this paper.

Inverted Detection and Control in Steering Vectors AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T15:00:38.184550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:00:38.184550Z digest=sha256:d69f8719a8f43501515f9700686f58f1e85bb725c871c584c2d7a7afd6a2eba3

Observation c2b9c053-81bb-4ef3-8471-255f8b679817 · inbound

ODRA: Synthesizing Cognitive Behavioral Therapy Sessions with Structured Chain-Of-Thought and Dynamic Patient Resistance cites this paper.

ODRA: Synthesizing Cognitive Behavioral Therapy Sessions with Structured Chain-Of-Thought and Dynamic Patient Resistance AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:06.016314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:56:06.016314Z digest=sha256:15c3477e036e6dd10b61776763c9823a9a58104132f4ff6c0956c1250eaa401a

Observation 52a1fa14-295f-46b4-87ae-5d6cc4258b0b · inbound

CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits cites this paper.

CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T00:27:20.702277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:27:20.702277Z digest=sha256:271a5826e6ded836513275e0939c793ec042ec914bf6e4b1edbb784d4bdc9e6a

Observation 3bc2e44e-d022-4b63-a352-5777f5fe446c · inbound

Scaling Inherently Interpretable Language Models cites this paper.

Scaling Inherently Interpretable Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:58.518062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:58.518062Z digest=sha256:d621d9390a1a4203e2f4ff8842cbeffbbb939b4985faaa0c4fbff523aae34d18