Pith. sign in

Paper Citation Record · LEDGER

SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2503.09532.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.09532 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T13:29:00.171371Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:29:45.394907Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 82a4d355-2fc2-4888-a10d-f5e945759d3e · inbound

Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy cites this paper.

Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.837701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.837701Z digest=sha256:5774d8e7a3f4b559f14e1a4418465d1a4c8d3a8da5bf66b9232fedd2e210635a

Observation 19903c6d-936d-43b0-9afe-7f2362cf3100 · inbound

Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures cites this paper.

Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:15.179663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:15.179663Z digest=sha256:a2ec4113fc6da8ef3586f77601d8ad3218664ac8e247521a111d14188c300236

Observation 620053fa-a4fe-44e0-a804-5c35444fc6ec · inbound

Resa: Transparent Reasoning Models via SAEs cites this paper.

Resa: Transparent Reasoning Models via SAEs SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:42.719642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:42.719642Z digest=sha256:089bb847a55a6c4d1c141e5e6e05fee0f60ff48c1ce0df395b068a47eeae975d

Observation 0a50bec3-c959-4b47-af7f-34a36a929314 · inbound

Evaluating SAE interpretability without explanations cites this paper.

Evaluating SAE interpretability without explanations SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.180244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:24:51.180244Z digest=sha256:00b16224c4abe94f5cac9b874c5b4751708d7a7d0ff49131dd17f1101a8655b8

Observation 4112d7ac-bf41-414e-88b3-355f5663fc65 · inbound

On the transferability of Sparse Autoencoders for interpreting compressed models cites this paper.

On the transferability of Sparse Autoencoders for interpreting compressed models SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:45.744009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:24:45.744009Z digest=sha256:36d5ae94f37c07cd2877e0d4c68889541396536136c924399213db1d726c8735

Observation 5c41650a-07f1-43c6-bec8-c9c5004543f8 · inbound

Distribution-Aware Feature Selection for SAEs cites this paper.

Distribution-Aware Feature Selection for SAEs SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:27:25.439860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:27:25.439860Z digest=sha256:a83c58fb3ebef9f4186de5fdfed5a277b2d9fda15596d36aac931fde347e6626

Observation 00bc17a8-ec58-433b-9b25-7291e4291f8c · inbound

Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework cites this paper.

Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:16:43.748207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-18T18:13:01.662828Z digest=sha256:3ff1bffc799cebfb270abde72f643e32837f2d317651dafbc6b90872d7e5816d

Observation eb0d2f34-a1c1-4baa-98c7-71091dd817b0 · inbound

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models cites this paper.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 148

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.605878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:e6baad982d37c842a2ce2264447cbd5e593f6c0d70e24c80e6e1a9f2dad539a2

Observation 1e7376d5-7454-4840-8813-48c7a2539620 · inbound

Stable and Steerable Sparse Autoencoders with Weight Regularization cites this paper.

Stable and Steerable Sparse Autoencoders with Weight Regularization SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T18:56:20.000242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:56:20.000242Z digest=sha256:564f07b204537bd2e8292df2984aa839c8e02abdc0fcb2bb652f823c45073f41

Observation f372237d-5f3a-434a-9a9c-253ec13b55e8 · inbound

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs cites this paper.

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:35:57.627704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T18:04:05.157103Z digest=sha256:fca4c0aff414c364946679a27882e06f160e2dcfebb73d055432c88ffea50740

Observation 41ce46e6-8ce3-437d-9d3d-838688b499cf · inbound

Structural Instability of Feature Composition cites this paper.

Structural Instability of Feature Composition SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:41:36.519790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T06:40:00.507484Z digest=sha256:2e6c72023d11a6e0ac879c89b185a13b9a8bc2a72f65ad184ab597d444747fb7

Observation fc8f724e-4403-4fbb-96fc-6f84e4825e25 · inbound

From Token Lists to Graph Motifs: Weisfeiler-Lehman Analysis of Sparse Autoencoder Features cites this paper.

From Token Lists to Graph Motifs: Weisfeiler-Lehman Analysis of Sparse Autoencoder Features SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:11.060148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T09:41:01.775116Z digest=sha256:3f3db8b5696ffa200e3e6e92cf9520883a8b3ce4cd0e0eae2bab7f690f5ae5ec

Observation 12da451f-70cb-4dca-a9f9-1ce293485063 · inbound

Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders cites this paper.

Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:15:54.303210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T03:13:58.543525Z digest=sha256:c7d91a245d1502d0081dd7a9bb0f64d5cfc150471d8fac317a00a7bce5b2197a

Observation 793d19d5-c2a0-4e38-a5d5-813223894e8a · inbound

Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders cites this paper.

Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:25.445401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T03:35:50.776347Z digest=sha256:8e8e74f6c9eff6933086db0786f510d93ebd47963fec4d647df1a8adda502c1d

Observation 47bdda45-8ee3-4658-8620-e9586d64adf1 · inbound

HH-SAE: Discovering and Steering Hierarchical Knowledge of Complex Manifolds cites this paper.

HH-SAE: Discovering and Steering Hierarchical Knowledge of Complex Manifolds SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:16:19.139898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T03:13:53.559096Z digest=sha256:6de219bd2135b3650e40f7a253c67b1a0f0707ad0dc4ce75b2c0d0ce84497e49

Observation 6ec363f6-219e-438e-a626-a06e5bd5fef3 · inbound

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations cites this paper.

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:23:30.730641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T14:16:44.232080Z digest=sha256:55d1e7a385cc68b683d4da81254d86e5a4cbcd80bd0eeac34ffb618489b5e74e

Observation a84392a0-db3d-4784-95e6-381b677202d3 · inbound

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations cites this paper.

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T05:02:51.380352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:02:51.380352Z digest=sha256:cbb0dd6484a147374c04044dd93dcc08f8a4c4e7f79b712c0673678ea0eef962

Observation 4091bcb3-81ce-41b3-8e70-c13ca7a8f2df · inbound

A Unifying Framework for Concept-Based Representational Similarity cites this paper.

A Unifying Framework for Concept-Based Representational Similarity SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:29.987745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T17:10:36.855674Z digest=sha256:4e62fcd5e6fe74d4221f9b11b48f211d54fd8ccbed4460de83e6eca8f68d14ed

Observation 2e6d96a4-bc02-44ab-b873-98d8a489405f · inbound

Do Sparse Autoencoders Learn Meaningful Concept Hierarchies? cites this paper.

Do Sparse Autoencoders Learn Meaningful Concept Hierarchies? SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:45.396552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T08:46:48.220801Z digest=sha256:68b05fa1e1abcc5f7815f20aae54219565bd06d17c5b69cc7bdf3b0b0ec447bd

Observation 054b1070-4bea-4917-a1ad-4f0284afe438 · inbound

Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models cites this paper.

Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T19:02:29.759226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T19:02:29.759226Z digest=sha256:6a67b838b31e33ce8ae320f00ca22dd3c0c3e5844e0bd6f0e665a05f4f34aa9c

Observation c7c1800e-f602-48f6-ab03-8968566103d0 · inbound

Decoder-Preserving Sparse Autoencoders: Which Readouts Survive Sparse Compression? cites this paper.

Decoder-Preserving Sparse Autoencoders: Which Readouts Survive Sparse Compression? SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:18.007688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:04:18.007688Z digest=sha256:31822476b4801ca491626d2f4c7a8feb46229b47b8881c24392a7aeb8a34e7f2

Observation b87c7a6c-e658-41fd-ac48-2b0833781747 · inbound

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects cites this paper.

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T10:03:56.301338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T10:03:56.301338Z digest=sha256:138f0d5c46761967b241ae8530a911d5b7853630c2a9a497dda01fb6223ed5e6

Observation 959c6a7f-12e1-46eb-8ee7-0a5a4c151dd9 · inbound

ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders cites this paper.

ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T07:58:32.246258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:58:32.246258Z digest=sha256:368c73c132a7e529692c15542c28ce4def3e12a8debf887319ba41493ee20d40

Observation cbbd8f16-e1d9-4f43-a074-7d7f632f70ea · inbound

Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation cites this paper.

Where You Measure Decides What You Measure: Position Selection in Ablation-Based SAE Evaluation SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T13:29:00.171371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:29:00.171371Z digest=sha256:fd0ac0b386c57c9c3f2dab7c4004a97caf1176a2774fa25cd20d7b7f696430de