Pith. sign in

Paper Citation Record · LEDGER

A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2502.17516.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.17516 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:17:55.972218Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9bc13ced-8fea-44d0-b73a-58019c5a9ef7 · inbound

LLM-Powered AI Agent Systems and Their Applications in Industry cites this paper.

LLM-Powered AI Agent Systems and Their Applications in Industry A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T14:06:37.994592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T14:05:54.535411Z digest=sha256:17fb7b5b4d1fc6040509d82d8407f70f85597926e41250ea46e5f3aefa0a69b8

Observation 7f7a930d-396f-4b96-8eb2-45c9931602cd · inbound

GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs cites this paper.

GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:08.637173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:45:08.637173Z digest=sha256:72b92757fbe79d9b67eb4bc029c6553ac8fb57a57f291a56f03f6eb21d1dc7d6

Observation c22e5c1e-ea0f-4fa7-b0fb-495a22d61f48 · inbound

V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models cites this paper.

V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:21:36.538648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-18T16:21:20.463222Z digest=sha256:9bd52822ea77d03b50144311e79c72685b2acb219004e21b216e6b1a0afa8ba2

Observation 8b8fabd1-c5e3-438d-995c-ed58af42c9ea · inbound

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models cites this paper.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 187

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:40:54.826576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:c692368252d90f9b7dfff5e384bbc3038319d6fa3bf530059046f712eeb31795

Observation c1a94952-6fb7-4e74-a193-88657279fc3f · inbound

Wearable AI in the Era of Large Sensor Models cites this paper.

Wearable AI in the Era of Large Sensor Models A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:30:58.075516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T16:35:36.541995Z digest=sha256:81252763ea9f8edbcf0228fc8348f42cd488d693ded87438a4ce6eed6fcddb32

Observation 80f20282-2381-4d37-ac40-15f948f5a11c · inbound

Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models cites this paper.

Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:11:53.748635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T07:02:02.752466Z digest=sha256:b05942b2b1f337aa5dca49529c3ca35cc78332c5988b55dca2f0da5827b7bc80

Observation 9c3af7c1-2df2-444b-ba12-b6abfec22626 · inbound

From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models cites this paper.

From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:46:10.141459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T05:44:37.891638Z digest=sha256:b7f3cba5c1f0cdaf758a595d0fe62e5d1e4ddfd9828da147cf8d329b2d6c3fa4

Observation f2855f91-1f3f-48db-b801-98d481d39e39 · inbound

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization cites this paper.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:28.897313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:ee49cae3a3e2fe32c9687578a2a541e2ab60639fc1cc24aa2ec99ee2bda930c9

Observation 3e6c05ac-1257-45cf-a960-2d6dd9c03c03 · inbound

Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models? cites this paper.

Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models? A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T06:31:10.545340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-22T06:26:19.580671Z digest=sha256:e845be45680b825d5c61d1349db1f3a3813c5a47599873520b25d77650f3eb72

Observation 783f212b-fc43-402d-a655-951a5df70ffb · inbound

The physics of AI weather models cites this paper.

The physics of AI weather models A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:25:14.433506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-25T02:20:39.109304Z digest=sha256:77993843fb5bd1e2064eb9c48ed88482c60c4b50c4615913e45f0c546fab1ca2

Observation e67e3ea1-366c-4208-a4c4-bd8dba2770c4 · inbound

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces cites this paper.

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-06-27T22:31:21.458106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T22:22:52.690010Z digest=sha256:e5a91550a3faacf7aa09d7932a171d4c8d018e60d112c634ca79a33801ea3e6f

Observation eb37bca8-edd0-43c5-9bb4-443e7ca2f79c · inbound

Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs cites this paper.

Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:39:24.638089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-26T19:36:40.901674Z digest=sha256:e633da2c4c65221c52006b04ecac2f659796287481d1b851d5a6e9efaf4872e4

Observation bdd90d0c-74b8-461e-995e-898a784fe6d3 · inbound

FairFlow: Demystifying and Mitigating Stereotype Bias in Text-to-Image Diffusion Transformers cites this paper.

FairFlow: Demystifying and Mitigating Stereotype Bias in Text-to-Image Diffusion Transformers A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T04:19:11.697328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T04:19:11.697328Z digest=sha256:3b9d662ed48d3c1621eeb9e250cacea6df2b319eae581fb6a3136dacf41d362f

Observation c4a06452-706b-4a01-866e-ced56ed66d42 · inbound

Unsupervised Features Mining via Activation Geometry cites this paper.

Unsupervised Features Mining via Activation Geometry A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T20:51:42.052454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:51:42.052454Z digest=sha256:0f0927ab8eae81375fb862d9d5b25e302e2c8418f16c283bf4b1274bee2f0873

Observation 5f34887d-bfc9-4ea3-84a8-1a5900ce8322 · inbound

How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA cites this paper.

How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T21:26:11.083317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:26:11.083317Z digest=sha256:0d9dea64179cc810fa8c8ce9d0da015ff7f437cac47470f3685369dc8fe7bc1b

Observation 3198cb64-e105-47b3-a1b1-53f7b87c65d9 · inbound

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs cites this paper.

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T01:33:51.250429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:33:51.250429Z digest=sha256:11e0903526312cc7f1b470e50b06635f4145281af8eb08381f387e6d06a8140b

Observation 5fabadee-bab1-4deb-b1f7-f996f485bf3d · inbound

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs cites this paper.

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:07.900572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:27:07.900572Z digest=sha256:dcef18a4d0b9a9d96519dd75a9c317fafdbcf7f5ccfa99ddddff3dd4db149016

Observation d6e0eebd-6527-463a-817f-2df09474b89b · inbound

Through the LENS: Local Geometric Decomposition of Vision-Language Model Representations cites this paper.

Through the LENS: Local Geometric Decomposition of Vision-Language Model Representations A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T00:46:24.026248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:46:24.026248Z digest=sha256:7ed77eb15b9243d9c2361447299cb12e3574786a3763bd09fa8c46f4e5db67dd

Observation 74cb08c5-3e38-4f66-b5a3-cf3e6d9061f1 · inbound

Multimodal Model Diffing for Feature Discovery and Control cites this paper.

Multimodal Model Diffing for Feature Discovery and Control A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T04:17:55.972218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:17:55.972218Z digest=sha256:e709f587b3aba02ca24d6581bcc7ab35fb655e3e21348deeb98c1dc31ad75be0