Pith. sign in

Paper Citation Record · LEDGER

Steering CLIP's vision transformer with sparse autoencoders

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2504.08729.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.08729 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:52:46.815838Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d35f7571-dfb2-4706-ad5b-ff8b8fc3d865 · inbound

Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering cites this paper.

Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering Steering CLIP's vision transformer with sparse autoencoders

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:37:15.823687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T11:35:01.836096Z digest=sha256:2435b4ad6f4e122864a18280160d57009b2844821a646697b2b08792ffbe5688

Observation 8344121f-33a7-4f0a-abfb-bbca9834bc8a · inbound

Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering cites this paper.

Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering Steering CLIP's vision transformer with sparse autoencoders

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:53:51.836956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:53:51.836956Z digest=sha256:85a473b188a736cf191c1f5cd4db87bcb2923bc6f77f42d630cd7cfe8d81ca75

Observation b4355137-defe-4ecc-9bf3-16a0739fbb7e · inbound

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders cites this paper.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Steering CLIP's vision transformer with sparse autoencoders

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:54742c1b9ec9132b9076e7be72fe32085fc7ffca7328433a189658f9d794e169

Observation 56aa0dc0-f95c-43e3-bd1e-933cb9b1b4bb · inbound

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs cites this paper.

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs Steering CLIP's vision transformer with sparse autoencoders

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:35:57.842361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:04:05.157103Z digest=sha256:a3376ef11f26b62fc424ae015f7265d6581560d5ce1e9d09426339fe93a948da

Observation 8bd5dcf7-0ce0-4bdb-a34e-98e8d13cf55b · inbound

From Attribution to Action: A Human-Centered Application of Activation Steering cites this paper.

From Attribution to Action: A Human-Centered Application of Activation Steering Steering CLIP's vision transformer with sparse autoencoders

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:21:04.771995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T14:59:12.966608Z digest=sha256:f8be5f6cd80c2e032ce85be3845e0ac6dd449438611f1383075930e4398ab46b

Observation bea5905e-54a0-4fb1-bd7f-39caf9eb5af5 · inbound

From Attribution to Action: A Human-Centered Application of Activation Steering cites this paper.

From Attribution to Action: A Human-Centered Application of Activation Steering Steering CLIP's vision transformer with sparse autoencoders

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T21:57:28.977088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:57:28.977088Z digest=sha256:4ae5eec410c03ef8484892325998c63ea41cc576558028e5bc6dc4a648067622

Observation 448f09ef-6c93-41b0-9206-d763d0df3f35 · inbound

LatentDiff: Scaling Semantic Dataset Comparison to Millions of Images cites this paper.

LatentDiff: Scaling Semantic Dataset Comparison to Millions of Images Steering CLIP's vision transformer with sparse autoencoders

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:21:06.381353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T20:16:30.677208Z digest=sha256:effb1208c811a2a9c1635738b902d00b1c430c69d2bb5447878b36e8865415ff

Observation b34ed1c4-f391-416a-8505-f9f43f1a1d9b · inbound

Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models cites this paper.

Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models Steering CLIP's vision transformer with sparse autoencoders

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:38:52.822726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T18:34:32.616965Z digest=sha256:4ba7f46a20fd94f2aace6653ff2450f521891af1aad75791eb4916d88ba56296

Observation e67791ac-c523-4e8d-84e7-fdc2638e14ed · inbound

TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment cites this paper.

TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment Steering CLIP's vision transformer with sparse autoencoders

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:47:10.093222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T22:22:01.979434Z digest=sha256:f3dd1216dfef796168790b33f85da96f105d3cfd4e3f06f2fc9d8449f354cdb9

Observation 60660a5b-47e4-4e5c-b7bb-3adff6a1f9ad · inbound

Vision-Language Asymmetry in Bistable Image Captioning cites this paper.

Vision-Language Asymmetry in Bistable Image Captioning Steering CLIP's vision transformer with sparse autoencoders

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:47:22.486788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T20:17:06.897180Z digest=sha256:cf27c7101926e3240011ee36f8116f0f4573722bda5d6cd0621264e13c41acbf

Observation a9661b87-f770-4b14-babd-9af400a18c38 · inbound

The Signs Were Always There: Training-Free Concept Detection and Steering in Raw Transformer Dimensions cites this paper.

The Signs Were Always There: Training-Free Concept Detection and Steering in Raw Transformer Dimensions Steering CLIP's vision transformer with sparse autoencoders

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T14:16:05.103480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:16:05.103480Z digest=sha256:ab82ba5e79ad0abc67c11bf3bfd72d47189f7aee0ff84fb4e126b7941768dd41

Observation 71047231-ddd8-4e20-95a6-74f4d2eba6cc · inbound

Can neurons speak? Semantic narration of vision at single-cell resolution cites this paper.

Can neurons speak? Semantic narration of vision at single-cell resolution Steering CLIP's vision transformer with sparse autoencoders

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:25.383766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T18:52:22.357765Z digest=sha256:f66caf14ace53fbe04851036e277d89de06192ed803f5bff5b7314392a38499f

Observation e9de1279-f54c-4fb6-bdef-a67a6bab760b · inbound

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization cites this paper.

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization Steering CLIP's vision transformer with sparse autoencoders

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-26T01:28:50.529147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T01:27:39.812228Z digest=sha256:2fc833d50285c5b3bb0ce9efcf6f52c0f62685221d87e8ba99db409bc02f9c21

Observation 97d1a87d-4e2c-4c1e-912e-7ad26715ad2e · inbound

IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers cites this paper.

IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers Steering CLIP's vision transformer with sparse autoencoders

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T04:50:36.937136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:50:36.937136Z digest=sha256:f63c3406944c4bf6770c14abd6d8cb713a2b98765136a8e6da41976f036a26af

Observation 2102290c-13d0-4391-a8e8-556bd4603a9d · inbound

IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers cites this paper.

IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers Steering CLIP's vision transformer with sparse autoencoders

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T16:52:46.815838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:52:46.815838Z digest=sha256:0732834f506169f6501e9c88e7aefbbba72b91e6723f1a6bcc6e3779de76175e