Pith. sign in

Paper Citation Record · LEDGER

Token Activation Map to Visually Explain Multimodal LLMs

As of 16 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 3 inbound Pith citation observations for arXiv:2506.23270.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23270 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:52:29.407583Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:58:38.554863Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T01:45:52.034775Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy47
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d4de39f2-8367-4a38-bd53-b02dca29f34c · outbound

This paper cites Quantifying attention flow in transformers.

Token Activation Map to Visually Explain Multimodal LLMs Quantifying attention flow in transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:39.315954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:22.828224Z digest=sha256:0db8c81c49ab979c3a640a806105e6ada7b7086fd8f6a94cd84057e38bc7526f

Observation ebc98010-930a-4a15-834f-258ab58ade04 · outbound

This paper cites Attnlrp: Attention- aware layer-wise relevance propagation for transformers.

Token Activation Map to Visually Explain Multimodal LLMs Attnlrp: Attention- aware layer-wise relevance propagation for transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:39.029975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:22.908507Z digest=sha256:b9cde4046b26841589fa80411be9bab34817f17a05ce512467fffb6bcaf83e1f

Observation bffa45d5-158c-4fba-819a-1a9fed0a37c3 · outbound

This paper cites Vl-interpret: An interactive visualization tool for interpreting vision-language transformers.

Token Activation Map to Visually Explain Multimodal LLMs Vl-interpret: An interactive visualization tool for interpreting vision-language transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:38.481689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:23.138905Z digest=sha256:959830c5e8d3188f733ba637ca17a8d82c1aae1f7981e9c97e156e3c7092851e

Observation 021db75b-a6aa-45f6-8a2f-df56eea62a9b · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Token Activation Map to Visually Explain Multimodal LLMs Flamingo: a visual language model for few-shot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.200165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.200165Z digest=sha256:3f1c90c716322eb4c64bc01207f63199e59cf96dd94b75242290e28b01c6e6aa

Observation e28f49e7-adc0-47b5-b63b-7b1ae3ee0216 · outbound

This paper cites Xai for trans- formers: Better explanations through conservative propa- gation.

Token Activation Map to Visually Explain Multimodal LLMs Xai for trans- formers: Better explanations through conservative propa- gation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:38.145550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:23.291980Z digest=sha256:5878e193d8c23ceefa7cf703071072195b49a299e7bc03a8ac46780f3629acf5

Observation 818ccc59-2c78-4235-9a49-8d9538a00479 · outbound

This paper cites Text2live: Text-driven layered image and video editing.

Token Activation Map to Visually Explain Multimodal LLMs Text2live: Text-driven layered image and video editing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:37.802003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:23.383073Z digest=sha256:1e389c00c53d1f839141b25500efbdddbc712b2b45a9fdc0a09c15b0ecde7eef

Observation 3aa7372b-2964-43cc-8255-0946c3531fbe · outbound

This paper cites Lvlm-intrepret: An interpretability tool for large vision-language models.

Token Activation Map to Visually Explain Multimodal LLMs Lvlm-intrepret: An interpretability tool for large vision-language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:37.419141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:23.464253Z digest=sha256:db144476aa3a096b2d1ae26599780cbf1dde5f7c1212e709f1000fd01dfb648f

Observation 56241116-b3c1-4dd0-826a-56f3dc54d316 · outbound

This paper cites An adaptive median filter for image denoising.

Token Activation Map to Visually Explain Multimodal LLMs An adaptive median filter for image denoising

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:37.138362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:23.579543Z digest=sha256:ef66c4852d59b9f395a9ebbefc1e7ae0a086ea42b0658b49f8eba3241b38c38d

Observation e807e267-b14c-43a4-9570-7995aaa5a0f0 · outbound

This paper cites Grad-cam++: General- ized gradient-based visual explanations for deep convolu- tional networks.

Token Activation Map to Visually Explain Multimodal LLMs Grad-cam++: General- ized gradient-based visual explanations for deep convolu- tional networks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.673939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.673939Z digest=sha256:01bb27836a553c67961740f17c39d51afccdb9cf039afd3f5f4e664fdb0a72ee

Observation 387ed1b8-79d1-4255-ad4f-234b5fe8ea23 · outbound

This paper cites Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers.

Token Activation Map to Visually Explain Multimodal LLMs Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.743484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.743484Z digest=sha256:704ddeefe0164b3dd775428edd57cae0139f899cb7087eed59e56df20601ac59

Observation f3551c4f-decd-459b-9d8c-447cd54b6c89 · outbound

This paper cites Transformer inter- pretability beyond attention visualization.

Token Activation Map to Visually Explain Multimodal LLMs Transformer inter- pretability beyond attention visualization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.945210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:23.851689Z digest=sha256:f988a6bf1d4a983fb44fbaa9ae5096640a91d056e7203d4ac4b33699372c667f

Observation b092b981-a4fe-4711-b9ed-7c5fca5c3b20 · outbound

This paper cites Less is more: Fewer interpretable region via submodular subset selection.

Token Activation Map to Visually Explain Multimodal LLMs Less is more: Fewer interpretable region via submodular subset selection

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.796077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:23.942557Z digest=sha256:4f743fdf101fb33f162d36fbdd2acb3a34a6aa626991dff35a876d4d43ea8ef7

Observation acc9bf12-e418-4552-8424-9ffd90af11f8 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Token Activation Map to Visually Explain Multimodal LLMs Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.065434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.065434Z digest=sha256:226b7a76cae404dbe2b5380fff844ba940f41289e3bec45c46d853957531014a

Observation afb0bd74-59d4-4c8c-af1c-87319c0c7084 · outbound

This paper cites Clip-ad: A language-guided staged dual-path model for zero-shot anomaly detection.

Token Activation Map to Visually Explain Multimodal LLMs Clip-ad: A language-guided staged dual-path model for zero-shot anomaly detection

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.619683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:24.182522Z digest=sha256:50116d15e63d8ea748a0db5bccf72039315fc2a41adb84307d3a6fa1e67278b8

Observation 60315eb3-3ff5-45aa-adee-6f7db657e89f · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Token Activation Map to Visually Explain Multimodal LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.270466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.270466Z digest=sha256:4a76c548368c3e4b9b8b0c38ecd7303daaab7dd15b8efa90dbd2e81e98ba4704

Observation 6b8b3449-02d5-490e-b819-49d285b05ae2 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Token Activation Map to Visually Explain Multimodal LLMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.447607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:24.355777Z digest=sha256:e0a8e05e91e3f9817fee685e75b4bad22d971bf7171b55e1026bdda5acfb4ac2

Observation d83c952b-6c0e-4d7f-9b64-d0d02c398361 · outbound

This paper cites A survey on multimodal large lan- guage models for autonomous driving.

Token Activation Map to Visually Explain Multimodal LLMs A survey on multimodal large lan- guage models for autonomous driving

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.265238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:24.418757Z digest=sha256:8223473131156109921de97a4ef21fd55cf7f08a9286d92c7e7bfd739ff03852

Observation 55e032c7-68f7-43e4-83e8-000f889c35f2 · outbound

This paper cites Flashattention: Fast and memory-efficient exact at- tention with io-awareness.

Token Activation Map to Visually Explain Multimodal LLMs Flashattention: Fast and memory-efficient exact at- tention with io-awareness

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.123765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:24.485401Z digest=sha256:3a4879e84a12136e409fbf5ca64d39407927b17a941a1f46952a5bd09980ad1e

Observation 562b6886-6ac5-4d32-9905-332e33589d3e · outbound

This paper cites Vision transformers need registers.

Token Activation Map to Visually Explain Multimodal LLMs Vision transformers need registers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.001289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:24.557594Z digest=sha256:9b45de6390e354d5305bdbff5024fcb77afc22e4c4ba6715b918f5a74f803571

Observation 9e0a0169-afc6-4f6d-9d91-4bc9b4618612 · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding.

Token Activation Map to Visually Explain Multimodal LLMs Bert: Pre-training of deep bidirectional trans- formers for language understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.628249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.628249Z digest=sha256:2094d7dad6de1fbae0b8d59f11a443acf722d8f6e199ad62c0e6111a94b5df83

Observation e29031a5-45de-4ba7-a704-248d703c7195 · outbound

This paper cites Hia: Towards chinese multimodal llms for comparative high-resolution joint diagnosis.

Token Activation Map to Visually Explain Multimodal LLMs Hia: Towards chinese multimodal llms for comparative high-resolution joint diagnosis

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.848320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:24.672164Z digest=sha256:8c6a58ca29c458026af4ef92dabcf9ec592ec50a4f91c988075b836035bd2baf

Observation 30eddce3-f46c-4af8-93d8-03533192b924 · outbound

This paper cites Holistic autonomous driving un- derstanding by bird’view injected multi-modal large models.

Token Activation Map to Visually Explain Multimodal LLMs Holistic autonomous driving un- derstanding by bird’view injected multi-modal large models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.696589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:24.721249Z digest=sha256:d77b9a9612507df3e0bf3f28c02a794f0e67b9e580488c116efb3849ebb6b6e1

Observation 06b09a51-da1c-4187-8bbb-77138383df40 · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

Token Activation Map to Visually Explain Multimodal LLMs An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.560451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:24.800458Z digest=sha256:f347a87fd4ce7c50fbeb590be9f538e7ca64984b9bff79855b4975e44a47b74a

Observation 0c3f78d9-6d9e-4956-ab1c-2079b9b30237 · outbound

This paper cites Deep residual learning for image recognition.

Token Activation Map to Visually Explain Multimodal LLMs Deep residual learning for image recognition

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.398468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:24.853813Z digest=sha256:301be8ca2c2011c3551a0c9300b8deac68282e0b0d9815579c73bf14eb3f49ae

Observation fceaabdb-ae1f-4e8c-a0e0-329b2a3b56d9 · outbound

This paper cites GPT-4o System Card.

Token Activation Map to Visually Explain Multimodal LLMs GPT-4o System Card

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.893909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.893909Z digest=sha256:679e97d2efcd18c344825e8477465a9f00635be64ca5d12082594a9447c6fc6c

Observation f3350316-b32e-4812-8ffb-e3e32ae89606 · outbound

This paper cites Layercam: Exploring hierarchical class activation maps for localization.

Token Activation Map to Visually Explain Multimodal LLMs Layercam: Exploring hierarchical class activation maps for localization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.237320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:24.962731Z digest=sha256:57c00730e0f3556c7076620c174fc5ac93b3d6ddebcddeef699a0a23224480f9

Observation 138f9eb9-5bc3-4bc4-98d9-eaf078fd15b0 · outbound

This paper cites Causal inference meets deep learning: A compre- hensive survey.

Token Activation Map to Visually Explain Multimodal LLMs Causal inference meets deep learning: A compre- hensive survey

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.053887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:25.019347Z digest=sha256:d2e2cb9a30d2343890196e74dc9029ecc1f1ebfbccff10165d00c018527dc88f

Observation 755369e1-7dd3-4e7d-ae27-6c9169819545 · outbound

This paper cites Unmasking clever hans predictors and as- sessing what machines really learn.

Token Activation Map to Visually Explain Multimodal LLMs Unmasking clever hans predictors and as- sessing what machines really learn

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.884300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:25.064591Z digest=sha256:bbf068446d4ddd42f240ef8cd5670e5ec4f573d03120118ce2b2836aea9c5a2f

Observation f96de966-94a9-40fb-a609-1d9f5bb8c4c6 · outbound

This paper cites Llava-docent: Instruction tuning with multimodal large lan- guage model to support art appreciation education.

Token Activation Map to Visually Explain Multimodal LLMs Llava-docent: Instruction tuning with multimodal large lan- guage model to support art appreciation education

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.688437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:25.118755Z digest=sha256:58cde677259c4f6dd6197b2f17b8102e549c51a7c24250fdcc97438288cbdd31

Observation 2da65a26-37a5-417a-ba21-7ff09e4f63ff · outbound

This paper cites Manipllm: Embodied multimodal large language model for object-centric robotic manipulation.

Token Activation Map to Visually Explain Multimodal LLMs Manipllm: Embodied multimodal large language model for object-centric robotic manipulation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.487383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:25.203550Z digest=sha256:099650692823300280e9b18e52e72117c0b7072005dbea7d2dd47a33a2b2b78e

Observation 591122c2-d792-48f2-9ef0-af9ae5158136 · outbound

This paper cites Exploring Visual Interpretability for Contrastive Language-Image Pre-training.

Token Activation Map to Visually Explain Multimodal LLMs Exploring Visual Interpretability for Contrastive Language-Image Pre-training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.254347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.254347Z digest=sha256:4d7ed69763f7dae305897f5eddd4d123afa1b78d692b9c82f2d0e44d8912892a

Observation 67b51a60-2297-48fb-a314-4198d1beea9a · outbound

This paper cites A closer look at the explainability of con- trastive language-image pre-training.

Token Activation Map to Visually Explain Multimodal LLMs A closer look at the explainability of con- trastive language-image pre-training

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.338754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:25.315356Z digest=sha256:ca5d00394893697d5b59f4ead5f74e30fd96967dcd2aeaf99d8846b4976ea0b1

Observation 0889dc48-50f3-439e-8bd1-2a2762db6951 · outbound

This paper cites Microsoft coco: Common objects in context.

Token Activation Map to Visually Explain Multimodal LLMs Microsoft coco: Common objects in context

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.381037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.381037Z digest=sha256:e69ed9e4d17f031d929b88032aef43f89b85520d4c67d987eb37ede1a8838f2c

Observation 8732b887-ceae-4a94-a807-544e0a93460e · outbound

This paper cites A medical multimodal large language model for future pandemics.

Token Activation Map to Visually Explain Multimodal LLMs A medical multimodal large language model for future pandemics

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.156408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:25.446880Z digest=sha256:c958004b52dfb97e021ca1608a6983385682242714bd1ec8bb616231ac21e01a

Observation 36fd000f-a697-4568-b524-23d0242845e7 · outbound

This paper cites Visual instruction tuning.

Token Activation Map to Visually Explain Multimodal LLMs Visual instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.016402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:25.514680Z digest=sha256:45fe38a37895a7259257db32a52e53b2b01763b9259814177053b3db73ec4fc4

Observation 3ed8483e-c879-4098-a0e1-200d1baa53ad · outbound

This paper cites A unified approach to interpreting model predictions.

Token Activation Map to Visually Explain Multimodal LLMs A unified approach to interpreting model predictions

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:33.817131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:25.583129Z digest=sha256:122cb57e7cb40df8f519b724ab81edc50b183eb0d284d9123ae302e93af65468

Observation 443a3d6b-8434-4fc5-9679-3d49fcda9473 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Token Activation Map to Visually Explain Multimodal LLMs Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:33.554791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:25.635669Z digest=sha256:ae75ac577dc340b18dfbc57a582811e32282c9fa2fd48fa64a5d913ac2eba9a7

Observation 01780bf3-78f0-4ac8-b323-3af92010a72c · outbound

This paper cites Smooth Grad-CAM++: An Enhanced Inference Level Visualization Technique for Deep Convolutional Neural Network Models.

Token Activation Map to Visually Explain Multimodal LLMs Smooth Grad-CAM++: An Enhanced Inference Level Visualization Technique for Deep Convolutional Neural Network Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.702588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.702588Z digest=sha256:8cc1d41981e42292b8ebc0825087ed716ad55c2196bc66868a33b8ad99f06675

Observation 2fdfa9fa-56af-4347-9d35-47813276b48a · outbound

This paper cites Towards vision-language mechanistic interpretabil- ity: A causal tracing tool for blip.

Token Activation Map to Visually Explain Multimodal LLMs Towards vision-language mechanistic interpretabil- ity: A causal tracing tool for blip

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:33.356930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:25.707466Z digest=sha256:44bc9160d4cb453fc76916f27bae246660e7cb63c1bdf6b9c21c48410a878d16

Observation 57c3a70d-810f-4ae9-8b7c-fb8d406672dc · outbound

This paper cites Causality.

Token Activation Map to Visually Explain Multimodal LLMs Causality

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.778705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.778705Z digest=sha256:30579939f740ce6f6f4e75ad9eb65d501e2e31f8f5a7115f4987607c7b01b6a4

Observation 5da96631-5200-416d-a009-0ce47338e1df · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Token Activation Map to Visually Explain Multimodal LLMs Learning transferable visual models from natural language supervi- sion

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:33.222645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:25.999812Z digest=sha256:d4f974c4c0a54c56882915c3ffa2ce5b9a7992210aacbef6a5600ed61ee87e69

Observation 6fae3910-a117-40af-a75a-4eb733df3128 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

Token Activation Map to Visually Explain Multimodal LLMs Glamm: Pixel grounding large multimodal model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.993373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:26.204832Z digest=sha256:fe90215f816c229ff37f7763a2f07d7c7e6bc56899401b7d0bbd371e4cccd499

Observation 5687849b-dc4f-402c-88d4-74696d16da54 · outbound

This paper cites ” why should i trust you?” explaining the predictions of any classifier.

Token Activation Map to Visually Explain Multimodal LLMs ” why should i trust you?” explaining the predictions of any classifier

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.802753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:26.369433Z digest=sha256:2a164134faa76aca82821535384be93479af85433fcd79c8a416b4cfcbe91d19

Observation 05a57b0f-e59c-410c-a9bb-cfb97897fd94 · outbound

This paper cites Causal interpretation of self-attention in pre-trained trans- formers.

Token Activation Map to Visually Explain Multimodal LLMs Causal interpretation of self-attention in pre-trained trans- formers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.626574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:26.613931Z digest=sha256:a4b5d4ec9666ec0b09d8a2b61009acfc4f6db865df12262fd2570c10e3d38304

Observation 196618f3-66a4-4c43-aee8-be7d529c1fed · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

Token Activation Map to Visually Explain Multimodal LLMs Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:26.849975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:26.849975Z digest=sha256:bca3a33d9914347bbb519024e5734a4d8c6febd9a75d07ff2fc4bbbe08b23897

Observation 67ed4cc3-904b-4d83-a56e-cd116cc1d3a2 · outbound

This paper cites Training- free object counting with prompts.

Token Activation Map to Visually Explain Multimodal LLMs Training- free object counting with prompts

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.256509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:27.179429Z digest=sha256:60913297ff7ef299f7ac922efd7e4686b7da2a41555eb5170724ebe4f771a66e

Observation 1a7d0272-266c-4853-b2cc-4f7fc734a287 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Token Activation Map to Visually Explain Multimodal LLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:27.282270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:27.282270Z digest=sha256:abfb4a64533f626693cd4bdc144dca3fcfee34b90626ecaf744b7b53536d8a31

Observation 6150cd2f-2db5-48e3-bbb1-95b8b2d7738d · outbound

This paper cites Understanding how vision-language models rea- son when solving visual math problems.

Token Activation Map to Visually Explain Multimodal LLMs Understanding how vision-language models rea- son when solving visual math problems

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.082031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:27.416113Z digest=sha256:94d839d2bd6a122ff8609b562bb7a919fd8cee5c26abf240f1084d9b4b7f63d7

Observation 7a765326-b36b-454a-9900-2259ac8c7c62 · outbound

This paper cites Attention is all you need.

Token Activation Map to Visually Explain Multimodal LLMs Attention is all you need

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.920473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:27.564729Z digest=sha256:4bb071c103caa88c6c7f4a2cda25a019a8558d83114e97a0835852bae7dcbe0b

Observation c69c4e1c-7ea3-477a-827b-b814833df753 · outbound

This paper cites Interpretable bilin- gual multimodal large language model for diverse biomed- ical tasks.

Token Activation Map to Visually Explain Multimodal LLMs Interpretable bilin- gual multimodal large language model for diverse biomed- ical tasks

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.726790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:27.740841Z digest=sha256:2d7aec9dc397b56ec0b11f499d8f7f25194f70c715569a32479c2231f252d6bd

Observation f3bf8d0d-6f75-4ecd-baf9-b4b3a078a944 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Token Activation Map to Visually Explain Multimodal LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:27.901257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:27.901257Z digest=sha256:2180ff0d6cc15bc68ae7aebe4c280acb8661e0a72fd12f6a0bcc66fe1212caf7

Observation 1593e6e8-1fc5-4776-a1c7-7620dcf83565 · outbound

This paper cites Star: A benchmark for situated reasoning in real-world videos.

Token Activation Map to Visually Explain Multimodal LLMs Star: A benchmark for situated reasoning in real-world videos

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.519512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:28.031705Z digest=sha256:c92cc068c465cc025d467ac4edded7191b787da8671aa758833bdfe44224ff1e

Observation 1e49da6f-40ab-42a7-aeb0-a8ae0eeb2f01 · outbound

This paper cites Efficient streaming language models with attention sinks.

Token Activation Map to Visually Explain Multimodal LLMs Efficient streaming language models with attention sinks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.306564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:28.184907Z digest=sha256:1bb5f8e4ae3d74af687317c6b442d0f05c9ea762eb61bc0ba42f3ba5c01b9bcd

Observation 84be4017-1d5c-4212-b382-2eb7d37008b0 · outbound

This paper cites A survey on causal inference.

Token Activation Map to Visually Explain Multimodal LLMs A survey on causal inference

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.119616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:28.331544Z digest=sha256:3c4cda1b1bedd89fc41531416ab6dda3402c4c62cc4542829a0a2a45be7131c7

Observation 0e52b13d-beb8-43b3-8e25-e2980e367942 · outbound

This paper cites From redundancy to relevance: Enhancing explainability in multimodal large language mod- els.

Token Activation Map to Visually Explain Multimodal LLMs From redundancy to relevance: Enhancing explainability in multimodal large language mod- els

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.950021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:28.453431Z digest=sha256:82d13ea4303468035bea03db8794a2f771dab3ba42b0661ae275e348d1e0adc1

Observation 0ec2ab28-29f7-4ab3-8a1a-d02842692370 · outbound

This paper cites Learning deep features for discrimina- tive localization.

Token Activation Map to Visually Explain Multimodal LLMs Learning deep features for discrimina- tive localization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:28.534320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:28.534320Z digest=sha256:6b318794239c3f48c5cffa780d6df72525003bdb362d84cdfadf65151f8263b1

Observation 9278089c-6546-41d2-b914-a25a9151eebb · outbound

This paper cites with” and the punctuation mark “.

Token Activation Map to Visually Explain Multimodal LLMs with” and the punctuation mark “

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.630735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:28.788264Z digest=sha256:d9188fbf67398fc0d2c5fca125c0678d2b6b9d9827c5aa425ff2dae7e3c595e8

Observation 6bf0e808-16c8-4475-b27e-9a9e1a53e8d0 · outbound

This paper cites an unresolved cited work.

Token Activation Map to Visually Explain Multimodal LLMs Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:30.447211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:28.941451Z digest=sha256:a99eeceb621523417cd4c5651ab8689503baebe8c166f9507984871445954f63

Observation 34a50edb-53d1-4e84-b268-dc2130cfdc34 · outbound

This paper cites living wall.

Token Activation Map to Visually Explain Multimodal LLMs living wall

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.265472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:29.101599Z digest=sha256:782140387b22a30b21df51afab2773fc2a50416afd323afce0128e021b6605e4

Observation cad3c0e3-9e47-4c87-b66d-b1722867a254 · outbound

This paper cites living wall.

Token Activation Map to Visually Explain Multimodal LLMs living wall

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.081243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:29.230235Z digest=sha256:8e01d22665bec5a716a25ccd91abbbd18946876a329cbaf34eaffa278000011a

Observation 0ec197a3-1b0d-48b1-85b4-bda3adb45db0 · outbound

This paper cites Object-determined.

Token Activation Map to Visually Explain Multimodal LLMs Object-determined

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:29.923294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:29.318102Z digest=sha256:27a56a3c9256fc2065c0d56cfe79598140ba2784c57935298a1a59a88778ea73

Observation 42bfa086-e8c6-463a-8d82-7eb6d2113a8c · outbound

This paper cites Missing arrows led to erroneous reasoning.

Token Activation Map to Visually Explain Multimodal LLMs Missing arrows led to erroneous reasoning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:29.760939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:29.407583Z digest=sha256:bba9f22caa312bde0f23bfcba5fc472dbc8c5da942887a6fe730cf6da803ba9e

Observation fa6da81a-05a1-4059-a00b-26a19f33bfe0 · outbound

This paper cites 2, 5, 6, 7, 14, 15.

Token Activation Map to Visually Explain Multimodal LLMs 2, 5, 6, 7, 14, 15

Reference 168

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:38.764893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:23.026379Z digest=sha256:043e59ae28535f82d978d53a5657ca857adbe4a6ea75a73c1d30dc50147fbfdc

Observation c8836885-99c1-429f-8b6c-71ada12ddd70 · outbound

This paper cites an unresolved cited work.

Token Activation Map to Visually Explain Multimodal LLMs Unresolved cited work

Reference 2016

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:30.757393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:28.664380Z digest=sha256:04fc556042f24267c8f95d7ecdc3cc92c0eb75c53e72a65965b4a9d2fed49be8

Observation ea323d48-56d5-4c18-97cd-36edf6b02d54 · outbound

This paper cites an unresolved cited work.

Token Activation Map to Visually Explain Multimodal LLMs Unresolved cited work

Reference 2017

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:32.418675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:52:27.014153Z digest=sha256:2477f5aeac7f069af2c15783b86f11d8d489296e859481dedb4e5b912ef0330e

Pith citing papers

Observation ee607a6c-424c-4103-8a9b-d5ac4adba379 · inbound

Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models cites this paper.

Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models Token Activation Map to Visually Explain Multimodal LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:38.554863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:58:38.554863Z digest=sha256:ca8e6cbf224f67b2f5eae3174f1917a1120b6f792b3074a307482611fcb5e076

Observation 61037376-a873-45fb-a438-086e6355c419 · inbound

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward cites this paper.

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward Token Activation Map to Visually Explain Multimodal LLMs

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:47.892061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T19:59:19.379119Z digest=sha256:9410676a767eda167b5c8f695ab2cfb2bfb85d8b8169478d6d700c41b9ede99a

Observation 68b4e094-aaed-4176-9567-dc7d2db5ea9d · inbound

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference cites this paper.

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference Token Activation Map to Visually Explain Multimodal LLMs

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:45:52.037170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-11T01:29:26.298131Z digest=sha256:25b1a1631e472812655407e9a114f800e244c6cee653f5fe27a9d3e6fec44e01