Pith. sign in

Paper Citation Record · LEDGER

Steering CLIP's vision transformer with sparse autoencoders

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2504.08729.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.08729 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:52:46.815838Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d35f7571-dfb2-4706-ad5b-ff8b8fc3d865 · inbound

Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering cites this paper.

Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering Steering CLIP's vision transformer with sparse autoencoders

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:37:15.823687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T11:35:01.836096Z digest=sha256:e4943a72cabba006faaa0ba2cc5744196a56ebe9e0165ace6cd48cfd28b7e7fa

Observation 8344121f-33a7-4f0a-abfb-bbca9834bc8a · inbound

Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering cites this paper.

Beyond Interpretability: When, Why, and How Sparse Autoencoders Enable Label-Free Visual Steering Steering CLIP's vision transformer with sparse autoencoders

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:53:51.836956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:53:51.836956Z digest=sha256:85a473b188a736cf191c1f5cd4db87bcb2923bc6f77f42d630cd7cfe8d81ca75

Observation b4355137-defe-4ecc-9bf3-16a0739fbb7e · inbound

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders cites this paper.

Interpreting Video Representations with Spatio-Temporal Sparse Autoencoders Steering CLIP's vision transformer with sparse autoencoders

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T12:01:02.997589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T12:01:02.997589Z digest=sha256:54742c1b9ec9132b9076e7be72fe32085fc7ffca7328433a189658f9d794e169

Observation 56aa0dc0-f95c-43e3-bd1e-933cb9b1b4bb · inbound

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs cites this paper.

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs Steering CLIP's vision transformer with sparse autoencoders

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:35:57.842361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:04:05.157103Z digest=sha256:33981955a10259b16dd6c435def53030e0702f3fdcfc554a78543e703a29682a

Observation 8bd5dcf7-0ce0-4bdb-a34e-98e8d13cf55b · inbound

From Attribution to Action: A Human-Centered Application of Activation Steering cites this paper.

From Attribution to Action: A Human-Centered Application of Activation Steering Steering CLIP's vision transformer with sparse autoencoders

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:21:04.771995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T14:59:12.966608Z digest=sha256:6ef6a277085fda1ded87aa5b9f047d5d8e5eb01b43cb74264f2847e0087a683d

Observation bea5905e-54a0-4fb1-bd7f-39caf9eb5af5 · inbound

From Attribution to Action: A Human-Centered Application of Activation Steering cites this paper.

From Attribution to Action: A Human-Centered Application of Activation Steering Steering CLIP's vision transformer with sparse autoencoders

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T21:57:28.977088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:57:28.977088Z digest=sha256:4ae5eec410c03ef8484892325998c63ea41cc576558028e5bc6dc4a648067622

Observation 448f09ef-6c93-41b0-9206-d763d0df3f35 · inbound

LatentDiff: Scaling Semantic Dataset Comparison to Millions of Images cites this paper.

LatentDiff: Scaling Semantic Dataset Comparison to Millions of Images Steering CLIP's vision transformer with sparse autoencoders

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:21:06.381353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T20:16:30.677208Z digest=sha256:a6dd3f480d6caae383922f0327f20dc3c1f1ba3bf22f5a3480f2756706d8a6c0

Observation b34ed1c4-f391-416a-8505-f9f43f1a1d9b · inbound

Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models cites this paper.

Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP models Steering CLIP's vision transformer with sparse autoencoders

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T18:38:52.822726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T18:34:32.616965Z digest=sha256:5320847786502bfd6508479360b585347066fa03c8adf20b3ad27d3283b25ad6

Observation e67791ac-c523-4e8d-84e7-fdc2638e14ed · inbound

TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment cites this paper.

TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment Steering CLIP's vision transformer with sparse autoencoders

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:47:10.093222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T22:22:01.979434Z digest=sha256:13b2bc67455aa079c80816f907076d35c819ce097d8d07cae8b9ed5874cdb917

Observation 60660a5b-47e4-4e5c-b7bb-3adff6a1f9ad · inbound

Vision-Language Asymmetry in Bistable Image Captioning cites this paper.

Vision-Language Asymmetry in Bistable Image Captioning Steering CLIP's vision transformer with sparse autoencoders

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T20:47:22.486788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T20:17:06.897180Z digest=sha256:5fd8da0d96b2b7e49d8a4121ebbeb61928c5534bc50943a2f43fa03764f9fa5e

Observation a9661b87-f770-4b14-babd-9af400a18c38 · inbound

The Signs Were Always There: Training-Free Concept Detection and Steering in Raw Transformer Dimensions cites this paper.

The Signs Were Always There: Training-Free Concept Detection and Steering in Raw Transformer Dimensions Steering CLIP's vision transformer with sparse autoencoders

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T14:16:05.103480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T14:16:05.103480Z digest=sha256:ab82ba5e79ad0abc67c11bf3bfd72d47189f7aee0ff84fb4e126b7941768dd41

Observation 71047231-ddd8-4e20-95a6-74f4d2eba6cc · inbound

Can neurons speak? Semantic narration of vision at single-cell resolution cites this paper.

Can neurons speak? Semantic narration of vision at single-cell resolution Steering CLIP's vision transformer with sparse autoencoders

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:25.383766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T18:52:22.357765Z digest=sha256:870b59261243ce8e45419ec13b0565b31aab45d76b1511812ab12e8a09376f93

Observation e9de1279-f54c-4fb6-bdef-a67a6bab760b · inbound

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization cites this paper.

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization Steering CLIP's vision transformer with sparse autoencoders

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-26T01:28:50.529147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T01:27:39.812228Z digest=sha256:77bac7dff483d0b8136b4dfc4c8a24fd43cd45d650fd6b83a0863700559f5801

Observation 97d1a87d-4e2c-4c1e-912e-7ad26715ad2e · inbound

IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers cites this paper.

IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers Steering CLIP's vision transformer with sparse autoencoders

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T04:50:36.937136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:50:36.937136Z digest=sha256:416c0d53306bf96a99a2520573d6fbfe8d21919f0b4a4f50d0764b62e7b48633

Observation 2102290c-13d0-4391-a8e8-556bd4603a9d · inbound

IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers cites this paper.

IRIS: A Visual Cortex-Inspired Framework for Analyzing Orientation Selectivity in Vision Transformers Steering CLIP's vision transformer with sparse autoencoders

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T16:52:46.815838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:52:46.815838Z digest=sha256:d9e394bac691b884ae08c78f64c35d6133227d827445eaa97d1520d0ce1cd873