Pith. sign in

Paper Citation Record · LEDGER

Interpreting CLIP's Image Representation via Text-Based Decomposition

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2310.05916.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.05916 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:51:30.398000Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T18:53:51.518085Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 269a60e8-f475-4424-ad0d-c0faa3b7516b · inbound

Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models cites this paper.

Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T13:15:10.673672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T13:15:10.632115Z digest=sha256:d569e769295a68d73949c670660b7c05de24bb8de491dc060bec090acc9b35dd

Observation defb6842-ee33-495f-bcea-0fa5f2286cda · inbound

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features cites this paper.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.398000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.398000Z digest=sha256:e4516c4df3f920e4466a6ba623421ece237d1c0dc6de309c39feb4ecd4a739d4

Observation d92a7e4f-2f38-449c-a7e6-46ee5da74063 · inbound

Understanding Design Fixation in Generative AI cites this paper.

Understanding Design Fixation in Generative AI Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T17:40:41.008598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:40:41.008598Z digest=sha256:09be8146a389d7d42f1e0a7dba415385e5574e6b76c1f89c06ed540904e4ca12

Observation 6835847c-c755-4c1d-af9a-2bbd3dec033f · inbound

Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads cites this paper.

Debiasing CLIP: Interpreting and Correcting Bias in Attention Heads Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:08.213795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:08.213795Z digest=sha256:0101c90574551dd9eb0e1103dfbf70ca33f9c3eea5704e7ab83dffc40135ed61

Observation 7a117c9c-7741-4049-a70a-582d506c319e · inbound

How Visual Representations Map to Language Feature Space in Multimodal LLMs cites this paper.

How Visual Representations Map to Language Feature Space in Multimodal LLMs Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:06.784761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:06:06.784761Z digest=sha256:2a1df2f17608977c4c18735ee24cb223fb9114da2afe5904973e4d03e68f9185

Observation d8934a82-4d82-429d-a342-57657f0fff5a · inbound

Model Science: getting serious about verification, explanation and control of AI systems cites this paper.

Model Science: getting serious about verification, explanation and control of AI systems Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T15:17:08.574138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:17:08.574138Z digest=sha256:b53a089f51035b87383bb30551feb5049f3a206c43d28c8018ecab820e339a35

Observation 3de5c3cf-6734-4c7e-9a89-c8653d1bf5e7 · inbound

CLIP-SVD: Efficient and Interpretable Vision-Language Adaptation via Singular Values cites this paper.

CLIP-SVD: Efficient and Interpretable Vision-Language Adaptation via Singular Values Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:56:45.986297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T18:55:31.923309Z digest=sha256:cf7e4d53791c6ae595a71a492c75f22c99737a17ca0a38a9a04126ddd19dacd6

Observation 070d1e97-0063-4c8a-a6a2-de68edb37979 · inbound

V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models cites this paper.

V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:21:36.519723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T16:21:20.463222Z digest=sha256:7f8d8640aeaf729c0b4a1ad33f7b601f424f40fdaaa3744046d46596fb8d391e

Observation 4d8f9fbf-e3fb-4b7d-b7dc-c047a79cca86 · inbound

The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail cites this paper.

The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T18:33:02.520085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:33:02.520085Z digest=sha256:ef02c3a84c49354bb6833e6a1e5db2fa287730993df50d02fd46a1463fd22f72

Observation a7f3c4ce-ed94-4a98-aa64-14e38daa0319 · inbound

BrainExplore: Large-Scale Discovery of Interpretable Visual Representations in the Human Brain cites this paper.

BrainExplore: Large-Scale Discovery of Interpretable Visual Representations in the Human Brain Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T17:41:51.180920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:41:51.180920Z digest=sha256:3ef5949381fab94caafe2fd6b9cfafc65490d2b35894347a7340fdcaec6460e9

Observation d0f8d83d-efdf-43ad-a2b8-496bcdca7b09 · inbound

MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence cites this paper.

MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T17:19:35.699200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:19:35.699200Z digest=sha256:c05835aca6f5ac91894bed748d5b3edb3020ab1c8b37fa65f84eb3e49f70b394

Observation 9d582a54-a184-477f-b3fd-d87d5e54a3d7 · inbound

Letting the neural code speak: Automated characterization of monkey visual neurons through human language cites this paper.

Letting the neural code speak: Automated characterization of monkey visual neurons through human language Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:12:07.046833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T02:12:00.335761Z digest=sha256:901c736e29677fb0d35a3818ec27243726da9080a81b274c643c2aa6be9f7025

Observation 027dbca0-7c2b-47ff-852d-dbbb4a3eaaa3 · inbound

Letting the neural code speak: Automated characterization of monkey visual neurons through human language cites this paper.

Letting the neural code speak: Automated characterization of monkey visual neurons through human language Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:19:02.997831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T21:17:53.615009Z digest=sha256:6f923f6396647c92d12b9d69ab101ecd4adcf0c31187cf92b3de026ed401dec3

Observation 52228df2-f26d-4141-a429-d6049a20b522 · inbound

AnimeAdapter: A Modular Adapter for Appearance-Consistent Anime Character Generation cites this paper.

AnimeAdapter: A Modular Adapter for Appearance-Consistent Anime Character Generation Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:44:53.085897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T08:44:51.204683Z digest=sha256:c64a651873aa86854fa4fb727736410e1dfa2482d6cca3ad5ba45b0c93760c5a

Observation a45e40e6-8243-4d25-8c3e-01f5315eb6a4 · inbound

AnchorDiff: Training-Free Concept Grounding for MM-DiTs via Anchor-Based Graph Propagation cites this paper.

AnchorDiff: Training-Free Concept Grounding for MM-DiTs via Anchor-Based Graph Propagation Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T18:53:51.520375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T18:48:47.286752Z digest=sha256:1a371b36c51715e5e37821827050de18091735f9f69003b4e0b7892363e3dba4