Pith. sign in

Paper Citation Record · LEDGER

A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 18 inbound Pith citation observations for arXiv:2502.17516.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.17516 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:45:08.637173Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9bc13ced-8fea-44d0-b73a-58019c5a9ef7 · inbound

LLM-Powered AI Agent Systems and Their Applications in Industry cites this paper.

LLM-Powered AI Agent Systems and Their Applications in Industry A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T14:06:37.994592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T14:05:54.535411Z digest=sha256:6bf2c52c9f02f015c4750172ada7520c42299d7f5bf0a4e8cb8db66dd715e0ba

Observation 7f7a930d-396f-4b96-8eb2-45c9931602cd · inbound

GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs cites this paper.

GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:08.637173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:45:08.637173Z digest=sha256:72b92757fbe79d9b67eb4bc029c6553ac8fb57a57f291a56f03f6eb21d1dc7d6

Observation c22e5c1e-ea0f-4fa7-b0fb-495a22d61f48 · inbound

V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models cites this paper.

V-SEAM: Visual Semantic Editing and Attention Modulating for Causal Interpretability of Vision-Language Models A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:21:36.538648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T16:21:20.463222Z digest=sha256:5ff83bf03c26852cae9d54a327423369daceb1bde125d09b756755f03c8b464b

Observation 8b8fabd1-c5e3-438d-995c-ed58af42c9ea · inbound

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models cites this paper.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 187

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:40:54.826576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:9944180e20d6cea9911ec2590f47dafd62bcdcc349bfa3f0205baf1df52729a4

Observation c1a94952-6fb7-4e74-a193-88657279fc3f · inbound

Wearable AI in the Era of Large Sensor Models cites this paper.

Wearable AI in the Era of Large Sensor Models A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:30:58.075516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:35:36.541995Z digest=sha256:333e3313dc46f8fe181bea2b861adf03ae55c64d34d465fa840ff9505235c452

Observation 80f20282-2381-4d37-ac40-15f948f5a11c · inbound

Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models cites this paper.

Beyond Text-Dominance: Understanding Modality Preference of Omni-modal Large Language Models A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:11:53.748635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T07:02:02.752466Z digest=sha256:64c540f6adb596a9109ca39fbdc46e9f75b6e685b10a4de658ceb9fd9e9dcbbe

Observation 9c3af7c1-2df2-444b-ba12-b6abfec22626 · inbound

From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models cites this paper.

From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:46:10.141459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T05:44:37.891638Z digest=sha256:630adeeceb825c4a2b2c698334ba20a4af78c8be65f553b1434976deda1258b1

Observation f2855f91-1f3f-48db-b801-98d481d39e39 · inbound

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization cites this paper.

SimReg: Achieving Higher Performance in the Pretraining via Embedding Similarity Regularization A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:16:28.897313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:33:45.192139Z digest=sha256:94c501cbf2706fa49ef69d55c81b73179bb49c02c2a10f6845f782046d2e8c0e

Observation 3e6c05ac-1257-45cf-a960-2d6dd9c03c03 · inbound

Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models? cites this paper.

Do Factual Recall Mechanisms Carry over from Text to Speech in Multimodal Language Models? A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T06:31:10.545340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T06:26:19.580671Z digest=sha256:d70ea3a85bca285072c5425c8f54b756dfa649fb3323dc8e6d92bad6507680dd

Observation 783f212b-fc43-402d-a655-951a5df70ffb · inbound

The physics of AI weather models cites this paper.

The physics of AI weather models A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:25:14.433506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-25T02:20:39.109304Z digest=sha256:08521750ce38c4dfd569fc1571a5fee594dd786325875b73e8fc422faf06b520

Observation e67e3ea1-366c-4208-a4c4-bd8dba2770c4 · inbound

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces cites this paper.

Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-06-27T22:31:21.458106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T22:22:52.690010Z digest=sha256:e5a9f2effd66caf61f48150f439350fe704982d528f96bbdd735ee366ffc47ed

Observation eb37bca8-edd0-43c5-9bb4-443e7ca2f79c · inbound

Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs cites this paper.

Who Wins the Conflict? Mechanistic Interpretability of Text Bias in Audio LLMs A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:39:24.638089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T19:36:40.901674Z digest=sha256:470cb0ffd1145e837643ab93d5dc2ab29d72267d178520367d395792fa90b9a4

Observation bdd90d0c-74b8-461e-995e-898a784fe6d3 · inbound

FairFlow: Demystifying and Mitigating Stereotype Bias in Text-to-Image Diffusion Transformers cites this paper.

FairFlow: Demystifying and Mitigating Stereotype Bias in Text-to-Image Diffusion Transformers A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T04:19:11.697328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T04:19:11.697328Z digest=sha256:3b9d662ed48d3c1621eeb9e250cacea6df2b319eae581fb6a3136dacf41d362f

Observation c4a06452-706b-4a01-866e-ced56ed66d42 · inbound

Unsupervised Features Mining via Activation Geometry cites this paper.

Unsupervised Features Mining via Activation Geometry A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-11T20:51:42.052454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T20:51:42.052454Z digest=sha256:0f0927ab8eae81375fb862d9d5b25e302e2c8418f16c283bf4b1274bee2f0873

Observation 5f34887d-bfc9-4ea3-84a8-1a5900ce8322 · inbound

How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA cites this paper.

How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T21:26:11.083317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:26:11.083317Z digest=sha256:0d9dea64179cc810fa8c8ce9d0da015ff7f437cac47470f3685369dc8fe7bc1b

Observation 3198cb64-e105-47b3-a1b1-53f7b87c65d9 · inbound

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs cites this paper.

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T01:33:51.250429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:33:51.250429Z digest=sha256:11e0903526312cc7f1b470e50b06635f4145281af8eb08381f387e6d06a8140b

Observation 5fabadee-bab1-4deb-b1f7-f996f485bf3d · inbound

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs cites this paper.

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:07.900572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:27:07.900572Z digest=sha256:dcef18a4d0b9a9d96519dd75a9c317fafdbcf7f5ccfa99ddddff3dd4db149016

Observation d6e0eebd-6527-463a-817f-2df09474b89b · inbound

Through the LENS: Local Geometric Decomposition of Vision-Language Model Representations cites this paper.

Through the LENS: Local Geometric Decomposition of Vision-Language Model Representations A Survey on Mechanistic Interpretability for Multi-Modal Foundation Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T00:46:24.026248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:46:24.026248Z digest=sha256:847d2ca4fb92c3d0f435cae6812e9ca0944ee4d3f3d900384ad3bf3cc4ee142a