Pith. sign in

Paper Citation Record · LEDGER

Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2412.02104.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02104 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T21:12:22.394230Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

10
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b9220300-1571-474a-9f81-64b400bdb17a · inbound

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning cites this paper.

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:32:32.888726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-23T04:30:38.804702Z digest=sha256:9c9f75381821f11b27e7b7d70a4eca48542c226ce5afe1a63dd57868f752beee

Observation f6ac4806-3d66-4213-bf96-9480f05b34b3 · inbound

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM cites this paper.

Survey on AI-Generated Media Detection: From Non-MLLM to MLLM Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T21:12:22.394230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:12:22.394230Z digest=sha256:118d07428d7fe0f60e563accb580d21cc353c570366ef1aa7ba0256936d6bbab

Observation a558fa4f-2a39-4149-80df-bd468a5a5042 · inbound

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring cites this paper.

Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T21:06:55.839319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:06:55.839319Z digest=sha256:80add8c2fcabddbf5d2b97f0907bdcb0679f8a4a9f521e4a4162d7e63e9ec67d

Observation 33d8aa99-94d9-4781-9d5c-7fddd82368c0 · inbound

CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring cites this paper.

CAFES: A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:32.345547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:32.345547Z digest=sha256:504cf5bbb713d97afd6c244b12e10297cc94aa054d8ea657b5fae128fc41b515

Observation dfe72256-4221-43c3-8a49-7ae6a8a501e9 · inbound

Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning cites this paper.

Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:19:44.885332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:19:44.885332Z digest=sha256:f25ecf821bf3a25e40ab5a1df1e41baf7af2d134fdbd2312d782dc7b42369792

Observation 58060430-2b71-43dc-95f0-8289fe674f1f · inbound

The time course of visuo-semantic representations in the human brain is captured by combining vision and language models cites this paper.

The time course of visuo-semantic representations in the human brain is captured by combining vision and language models Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:33.412140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:33.412140Z digest=sha256:c545a0154db642650f7bb0631edb27b9f4a97e60b3f71e7e92aed988943cdf02

Observation 139b1daa-47d7-46ab-8b81-4a9e375eafdd · inbound

Towards Transparent AI: A Survey on Explainable Large Language Models cites this paper.

Towards Transparent AI: A Survey on Explainable Large Language Models Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:42.220033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:42.220033Z digest=sha256:e612830be634ba215a0095752392c859d33ea0c7e2c64800ae86ee776b5392f3

Observation 6f02e5a7-a4d0-4819-a23e-96ab81ce1ba8 · inbound

LongAnimation: Long Animation Generation with Dynamic Global-Local Memory cites this paper.

LongAnimation: Long Animation Generation with Dynamic Global-Local Memory Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:43:59.901023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:43:59.901023Z digest=sha256:be60ad5ab6d88e6c02768c7c72d6d6834d80471efa7acf813a42259ebdc43934

Observation 1a94cb14-3b90-4d4a-86b4-5b968eb7dec5 · inbound

Recourse, Repair, Reparation, & Prevention: A Stakeholder Analysis of AI Supply Chains cites this paper.

Recourse, Repair, Reparation, & Prevention: A Stakeholder Analysis of AI Supply Chains Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:59.826071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:59.826071Z digest=sha256:798204384be0b88a6cc010c3cb1cecdb646c77a1ff8187b1f5c97e13614807f3

Observation 91c07260-d94d-42fb-9594-9506434b01f1 · inbound

Benchmarking Foundation Models with Multimodal Public Electronic Health Records cites this paper.

Benchmarking Foundation Models with Multimodal Public Electronic Health Records Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T15:50:15.335365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:50:15.335365Z digest=sha256:bbe6f8e4120b12263c3e8f8f9f23e707832b50caafe0fbf54198fbd5cbc778f4

Observation a855dbf3-1936-4255-8ffb-2b16a7239d31 · inbound

SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law cites this paper.

SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:37:19.033109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:37:19.033109Z digest=sha256:8f89ab5e5f046328b2683902a7ca6bafa87df99d7c22b642dcbaec43e0c4ad7f

Observation 67fde6f4-d426-4769-bb32-61f74d561522 · inbound

Multimodal Function Vectors for Visual Relations cites this paper.

Multimodal Function Vectors for Visual Relations Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T12:44:41.600573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:44:41.600573Z digest=sha256:5f63c6024a7690923f978e10af42f9b2e99b98884757891ed41bb875005fa716

Observation bb2e2c30-9dc5-43f9-a2d5-8101406fb5b0 · inbound

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework cites this paper.

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T20:19:13.457923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:19:13.457923Z digest=sha256:36ffa0f45f9185b6eb4aa9d678c72e508cd5e8d09cf6e678df36441f8b7ad8b9

Observation 40dc1e0e-3e8d-458a-ab18-7b9aad361e32 · inbound

Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions cites this paper.

Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:51:30.385816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T18:50:52.313363Z digest=sha256:fa47eef3fe2ef977f5ee32857d3236f3f135287a5d880dd07323742a35caf699

Observation b8fdaef4-9574-46bd-9444-b1ee9cac28f8 · inbound

Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions cites this paper.

Beyond Explainable AI (XAI): An Overdue Paradigm Shift and Post-XAI Research Directions Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T20:04:32.769786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:04:32.769786Z digest=sha256:3f09ffeb8f62490c0794125fc4511bacad4265540357e0481b343b9be6a1d87e

Observation 53edac71-202e-4cd9-b8b8-37c348243d79 · inbound

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks cites this paper.

Explicit Logic Channel for Validation and Enhancement of MLLMs on Zero-Shot Tasks Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:10:01.961487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T11:09:51.816554Z digest=sha256:95c575c8e596bc8122fb6c2fb0f3cec6aef3cc35160da394bc9e56430a67fe96

Observation 9ada7a52-d94c-42ab-8348-381645d93efc · inbound

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing cites this paper.

UHR-BAT: Budget-Aware Token Compression Vision-Language model for Ultra-High-Resolution Remote Sensing Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:55:28.864017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:53:13.255412Z digest=sha256:0eb046fd50a234e58abb67eab8d9eb4b3e2c037e99dfb6c24e82972dd91f8047

Observation f4f1e593-6c83-4209-b8a9-127fd4fdbe67 · inbound

From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models cites this paper.

From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:46:09.455482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T05:44:37.891638Z digest=sha256:3a762acda8223dcfe20b008822eeff07ecd496822966b8e236ebb364268620e3

Observation 85b28b53-d8f1-4a66-8350-cfa4c018e985 · inbound

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety cites this paper.

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:05.779290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T03:00:34.862711Z digest=sha256:a1aa32062cd88554b166590e83e04f0211b4177b320afca7a80902720eb1984c

Observation 15f7d97c-7a62-4810-a9c1-7923d25e4cb0 · inbound

3D-CBM: A Framework for Concept-Based Interpretability in Generative 3D Modeling cites this paper.

3D-CBM: A Framework for Concept-Based Interpretability in Generative 3D Modeling Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:27:40.524932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T13:10:42.454041Z digest=sha256:76d6531d38733653b5fb90a0adfd7129728349645406da1ed71f8c768b27784e

Observation 4ea7972b-7e10-44d2-ab6a-ca94eda19b02 · inbound

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs cites this paper.

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T01:33:50.503472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:33:50.503472Z digest=sha256:a5b295115d1ecf204a326e7c90fc7a7fc999e6cd126ae3c0f4cb6634b248306d

Observation 763be5cb-b83e-42dd-b839-0eec196a34bc · inbound

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs cites this paper.

Which Modality Decides? Counterfactual Modality Attribution for Multimodal LLMs Explainable and Interpretable Multimodal Large Language Models: A Comprehensive Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T04:27:07.859701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:27:07.859701Z digest=sha256:930596bffcd8f9e291fefc7e2dfd87e59077012e0cbcf742f90d609fedae50fa