Pith. sign in

Paper Citation Record · LEDGER

Token Activation Map to Visually Explain Multimodal LLMs

As of 9 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 3 inbound Pith citation observations for arXiv:2506.23270.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23270 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:52:29.407583Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:58:38.554863Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T01:45:52.034775Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy47
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d4de39f2-8367-4a38-bd53-b02dca29f34c · outbound

This paper cites Quantifying attention flow in transformers.

Token Activation Map to Visually Explain Multimodal LLMs Quantifying attention flow in transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:39.315954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:22.828224Z digest=sha256:72d6f77ca783fe82caab0945c84fa9dd83d37095edff45b39ba2cf89bc11574f

Observation ebc98010-930a-4a15-834f-258ab58ade04 · outbound

This paper cites Attnlrp: Attention- aware layer-wise relevance propagation for transformers.

Token Activation Map to Visually Explain Multimodal LLMs Attnlrp: Attention- aware layer-wise relevance propagation for transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:39.029975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:22.908507Z digest=sha256:71251925dbc151d8dbce457542e2e12098890a8d6ebf515bee65de1b236f62a3

Observation bffa45d5-158c-4fba-819a-1a9fed0a37c3 · outbound

This paper cites Vl-interpret: An interactive visualization tool for interpreting vision-language transformers.

Token Activation Map to Visually Explain Multimodal LLMs Vl-interpret: An interactive visualization tool for interpreting vision-language transformers

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:38.481689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:23.138905Z digest=sha256:d4849948af20c59536d5cf5669ea0c27d60dfb688fa0d66c12cd8c08ba0dab53

Observation 021db75b-a6aa-45f6-8a2f-df56eea62a9b · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Token Activation Map to Visually Explain Multimodal LLMs Flamingo: a visual language model for few-shot learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.200165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.200165Z digest=sha256:9b09623f06d6d6bb2fbf9bc258d1867d1c0960d07e548671287d3c396c6d87b3

Observation e28f49e7-adc0-47b5-b63b-7b1ae3ee0216 · outbound

This paper cites Xai for trans- formers: Better explanations through conservative propa- gation.

Token Activation Map to Visually Explain Multimodal LLMs Xai for trans- formers: Better explanations through conservative propa- gation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:38.145550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:23.291980Z digest=sha256:b7dc65321d25d58026fc4aa6eebee99741e02084b9211aafef0e07d084d31b94

Observation 818ccc59-2c78-4235-9a49-8d9538a00479 · outbound

This paper cites Text2live: Text-driven layered image and video editing.

Token Activation Map to Visually Explain Multimodal LLMs Text2live: Text-driven layered image and video editing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:37.802003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:23.383073Z digest=sha256:f099ae4036b925068090b7d2fe2963ddc9b1d8ea1ad29de1de9c317ec309e90b

Observation 3aa7372b-2964-43cc-8255-0946c3531fbe · outbound

This paper cites Lvlm-intrepret: An interpretability tool for large vision-language models.

Token Activation Map to Visually Explain Multimodal LLMs Lvlm-intrepret: An interpretability tool for large vision-language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:37.419141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:23.464253Z digest=sha256:7299fce20cec3bebe535cd87d9c448e8aa236b848f310b2c45da21fb73ef7524

Observation 56241116-b3c1-4dd0-826a-56f3dc54d316 · outbound

This paper cites An adaptive median filter for image denoising.

Token Activation Map to Visually Explain Multimodal LLMs An adaptive median filter for image denoising

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:37.138362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:23.579543Z digest=sha256:4b1195fe954f9fd021942ba6ec04a6a3cd5440085c9e4043458c820a3f55be15

Observation e807e267-b14c-43a4-9570-7995aaa5a0f0 · outbound

This paper cites Grad-cam++: General- ized gradient-based visual explanations for deep convolu- tional networks.

Token Activation Map to Visually Explain Multimodal LLMs Grad-cam++: General- ized gradient-based visual explanations for deep convolu- tional networks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.673939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.673939Z digest=sha256:85c9f5f7c7280fa798a38d9f66ccc6fb726d8a9df92eb1cd251c26ec1d106dd2

Observation 387ed1b8-79d1-4255-ad4f-234b5fe8ea23 · outbound

This paper cites Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers.

Token Activation Map to Visually Explain Multimodal LLMs Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:23.743484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:23.743484Z digest=sha256:f1dbeebf02accb0a53f473c043b91c3df81a3c9bd7f69e9e0b20b6b178b12435

Observation f3551c4f-decd-459b-9d8c-447cd54b6c89 · outbound

This paper cites Transformer inter- pretability beyond attention visualization.

Token Activation Map to Visually Explain Multimodal LLMs Transformer inter- pretability beyond attention visualization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.945210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:23.851689Z digest=sha256:5cbd53745cfddd504f2311be3ef268f4a29d36b4c1dbef4183f3907b83d246b6

Observation b092b981-a4fe-4711-b9ed-7c5fca5c3b20 · outbound

This paper cites Less is more: Fewer interpretable region via submodular subset selection.

Token Activation Map to Visually Explain Multimodal LLMs Less is more: Fewer interpretable region via submodular subset selection

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.796077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:23.942557Z digest=sha256:a0a50ae362621c678ebadd83a91a625dee7c5cef3f931c22d1220495475b671d

Observation acc9bf12-e418-4552-8424-9ffd90af11f8 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Token Activation Map to Visually Explain Multimodal LLMs Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.065434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.065434Z digest=sha256:384038f6ac417f47aa323854163531b0b1ece14716243222bfa83cd2e2123dc1

Observation afb0bd74-59d4-4c8c-af1c-87319c0c7084 · outbound

This paper cites Clip-ad: A language-guided staged dual-path model for zero-shot anomaly detection.

Token Activation Map to Visually Explain Multimodal LLMs Clip-ad: A language-guided staged dual-path model for zero-shot anomaly detection

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.619683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:24.182522Z digest=sha256:5fd775c8ea148aae034dc51b980a1de3c1159115deeee286f251dfd1e39a104b

Observation 60315eb3-3ff5-45aa-adee-6f7db657e89f · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Token Activation Map to Visually Explain Multimodal LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.270466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.270466Z digest=sha256:131914aaa72c23693595b03bfd690a1bff82382e30ece3b7faa5f52c71bf44b7

Observation 6b8b3449-02d5-490e-b819-49d285b05ae2 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Token Activation Map to Visually Explain Multimodal LLMs Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.447607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:24.355777Z digest=sha256:42c9f91e06edabb8ca22187eb62b2660b97f0ac17ff15e6365419bd451fff8be

Observation d83c952b-6c0e-4d7f-9b64-d0d02c398361 · outbound

This paper cites A survey on multimodal large lan- guage models for autonomous driving.

Token Activation Map to Visually Explain Multimodal LLMs A survey on multimodal large lan- guage models for autonomous driving

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.265238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:24.418757Z digest=sha256:cd0128c574f9842af32aa40e47aca261074670e91bb446dfc29285586efca459

Observation 55e032c7-68f7-43e4-83e8-000f889c35f2 · outbound

This paper cites Flashattention: Fast and memory-efficient exact at- tention with io-awareness.

Token Activation Map to Visually Explain Multimodal LLMs Flashattention: Fast and memory-efficient exact at- tention with io-awareness

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.123765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:24.485401Z digest=sha256:c891c18df2bb44899f05ae3e06d4e900abc542a2e1b95bbdd5277ba713329d72

Observation 562b6886-6ac5-4d32-9905-332e33589d3e · outbound

This paper cites Vision transformers need registers.

Token Activation Map to Visually Explain Multimodal LLMs Vision transformers need registers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:36.001289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:24.557594Z digest=sha256:010dbac0343d0bea128a3e269bfb9c3e4aca098668f1df36f826b9b98201bd1f

Observation 9e0a0169-afc6-4f6d-9d91-4bc9b4618612 · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding.

Token Activation Map to Visually Explain Multimodal LLMs Bert: Pre-training of deep bidirectional trans- formers for language understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.628249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.628249Z digest=sha256:5dfcd8769c1e1d8ca0760897fb3163e6cfe9ca1ec93337ae52a68fceb89bacf9

Observation e29031a5-45de-4ba7-a704-248d703c7195 · outbound

This paper cites Hia: Towards chinese multimodal llms for comparative high-resolution joint diagnosis.

Token Activation Map to Visually Explain Multimodal LLMs Hia: Towards chinese multimodal llms for comparative high-resolution joint diagnosis

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.848320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:24.672164Z digest=sha256:9ca182546df6203cfc6cd9524c4e8835121cd72f23570928f82caa5296b3aa03

Observation 30eddce3-f46c-4af8-93d8-03533192b924 · outbound

This paper cites Holistic autonomous driving un- derstanding by bird’view injected multi-modal large models.

Token Activation Map to Visually Explain Multimodal LLMs Holistic autonomous driving un- derstanding by bird’view injected multi-modal large models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.696589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:24.721249Z digest=sha256:48dc2c7b3cc1ec95335db8d2943c86dde8757e03d4aeb6673c843b8e38ecff63

Observation 06b09a51-da1c-4187-8bbb-77138383df40 · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

Token Activation Map to Visually Explain Multimodal LLMs An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.560451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:24.800458Z digest=sha256:da00cf8ac28f0d79101f041c7fdb4c54ece5b673aa9ab1e3748021fcee5db815

Observation 0c3f78d9-6d9e-4956-ab1c-2079b9b30237 · outbound

This paper cites Deep residual learning for image recognition.

Token Activation Map to Visually Explain Multimodal LLMs Deep residual learning for image recognition

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.398468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:24.853813Z digest=sha256:45c7f7e5e7bb21fb1bf81819ee6c9b1651a59ff753ea2b4911dbf3c1170810d2

Observation fceaabdb-ae1f-4e8c-a0e0-329b2a3b56d9 · outbound

This paper cites GPT-4o System Card.

Token Activation Map to Visually Explain Multimodal LLMs GPT-4o System Card

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:24.893909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:24.893909Z digest=sha256:3575e5d8a15a01e420aca9e311a91dd3f1456f59e3f703794adc2786a6cbd83c

Observation f3350316-b32e-4812-8ffb-e3e32ae89606 · outbound

This paper cites Layercam: Exploring hierarchical class activation maps for localization.

Token Activation Map to Visually Explain Multimodal LLMs Layercam: Exploring hierarchical class activation maps for localization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.237320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:24.962731Z digest=sha256:67034f7ad0654071b16f8fe964498b7d41d95dbb214f3abf32cfc0775c2f1b3d

Observation 138f9eb9-5bc3-4bc4-98d9-eaf078fd15b0 · outbound

This paper cites Causal inference meets deep learning: A compre- hensive survey.

Token Activation Map to Visually Explain Multimodal LLMs Causal inference meets deep learning: A compre- hensive survey

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:35.053887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:25.019347Z digest=sha256:bf268bef6de8fda0c79b02d1240f847e9b2d993d4bc6c3ef2e5b3ef4ec02bb04

Observation 755369e1-7dd3-4e7d-ae27-6c9169819545 · outbound

This paper cites Unmasking clever hans predictors and as- sessing what machines really learn.

Token Activation Map to Visually Explain Multimodal LLMs Unmasking clever hans predictors and as- sessing what machines really learn

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.884300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:25.064591Z digest=sha256:b0f0e73e14381304a8f215395cbb843f5bf5f128efb027fd4244686cb52c621d

Observation f96de966-94a9-40fb-a609-1d9f5bb8c4c6 · outbound

This paper cites Llava-docent: Instruction tuning with multimodal large lan- guage model to support art appreciation education.

Token Activation Map to Visually Explain Multimodal LLMs Llava-docent: Instruction tuning with multimodal large lan- guage model to support art appreciation education

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.688437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:25.118755Z digest=sha256:42c3851c8cbe9ef92045cf8de96b22a60688149116d9e0ba93c26353d3823611

Observation 2da65a26-37a5-417a-ba21-7ff09e4f63ff · outbound

This paper cites Manipllm: Embodied multimodal large language model for object-centric robotic manipulation.

Token Activation Map to Visually Explain Multimodal LLMs Manipllm: Embodied multimodal large language model for object-centric robotic manipulation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.487383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:25.203550Z digest=sha256:5e7411c9597f38555177712cccba0e42e841f8b6fe2a62113e590c67eac47c0b

Observation 591122c2-d792-48f2-9ef0-af9ae5158136 · outbound

This paper cites Exploring Visual Interpretability for Contrastive Language-Image Pre-training.

Token Activation Map to Visually Explain Multimodal LLMs Exploring Visual Interpretability for Contrastive Language-Image Pre-training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.254347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.254347Z digest=sha256:3346f2e8c38f46c5d725e7ff9880b2309d4a6e14dab4a0ff520e58508267fb8f

Observation 67b51a60-2297-48fb-a314-4198d1beea9a · outbound

This paper cites A closer look at the explainability of con- trastive language-image pre-training.

Token Activation Map to Visually Explain Multimodal LLMs A closer look at the explainability of con- trastive language-image pre-training

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.338754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:25.315356Z digest=sha256:20cda67d60a0cf457a5bb16dbad7f573a9b805e6b3e1dba52fde59e3ac0cb8ed

Observation 0889dc48-50f3-439e-8bd1-2a2762db6951 · outbound

This paper cites Microsoft coco: Common objects in context.

Token Activation Map to Visually Explain Multimodal LLMs Microsoft coco: Common objects in context

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.381037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.381037Z digest=sha256:6c71e35b5a753e3c81eb296da0877cca2b70289af88081b94966209b61c49382

Observation 8732b887-ceae-4a94-a807-544e0a93460e · outbound

This paper cites A medical multimodal large language model for future pandemics.

Token Activation Map to Visually Explain Multimodal LLMs A medical multimodal large language model for future pandemics

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.156408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:25.446880Z digest=sha256:ca9001abfeb1d426c7a5c1af58166d8233995f31c2e2c75fdac6b74b1d54a9e9

Observation 36fd000f-a697-4568-b524-23d0242845e7 · outbound

This paper cites Visual instruction tuning.

Token Activation Map to Visually Explain Multimodal LLMs Visual instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:34.016402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:25.514680Z digest=sha256:7eaf7d33592773d760c36768dc87ec926ea864b17ffcdd59eb26a395cedf196d

Observation 3ed8483e-c879-4098-a0e1-200d1baa53ad · outbound

This paper cites A unified approach to interpreting model predictions.

Token Activation Map to Visually Explain Multimodal LLMs A unified approach to interpreting model predictions

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:33.817131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:25.583129Z digest=sha256:b4ef05fcc463073b51072709afe75e7f3fd8ad846710b4661f00efcb66c6cdd5

Observation 443a3d6b-8434-4fc5-9679-3d49fcda9473 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Token Activation Map to Visually Explain Multimodal LLMs Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:33.554791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:25.635669Z digest=sha256:234978d9489e19630b4e5d8fa5a5ae7daf98abfe5bf81b54f457e9502b6b495e

Observation 01780bf3-78f0-4ac8-b323-3af92010a72c · outbound

This paper cites Smooth Grad-CAM++: An Enhanced Inference Level Visualization Technique for Deep Convolutional Neural Network Models.

Token Activation Map to Visually Explain Multimodal LLMs Smooth Grad-CAM++: An Enhanced Inference Level Visualization Technique for Deep Convolutional Neural Network Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.702588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.702588Z digest=sha256:7a8a56a075dcf71b7e1c0167a1c307e6e07dda3e63de58e743e432d7416c4cda

Observation 2fdfa9fa-56af-4347-9d35-47813276b48a · outbound

This paper cites Towards vision-language mechanistic interpretabil- ity: A causal tracing tool for blip.

Token Activation Map to Visually Explain Multimodal LLMs Towards vision-language mechanistic interpretabil- ity: A causal tracing tool for blip

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:33.356930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:25.707466Z digest=sha256:872f043dfbcd65b77715abc29d61b8f6bb2950aae9c438f80bdb00884a955cdd

Observation 57c3a70d-810f-4ae9-8b7c-fb8d406672dc · outbound

This paper cites Causality.

Token Activation Map to Visually Explain Multimodal LLMs Causality

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:25.778705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:25.778705Z digest=sha256:d85bf057ab30cbf7bd0b7b67413486aa4aaf2e6c2fbffcaac3e774bd29846b79

Observation 5da96631-5200-416d-a009-0ce47338e1df · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Token Activation Map to Visually Explain Multimodal LLMs Learning transferable visual models from natural language supervi- sion

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:33.222645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:25.999812Z digest=sha256:3119747064eb582c65f5bc218fa31826b7331ffa3edfae528fe5ede0f2a27d64

Observation 6fae3910-a117-40af-a75a-4eb733df3128 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

Token Activation Map to Visually Explain Multimodal LLMs Glamm: Pixel grounding large multimodal model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.993373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:26.204832Z digest=sha256:69bd5117d84a67ccfbd120d267108d2c12187594081231d920f9225c7b2bf26a

Observation 5687849b-dc4f-402c-88d4-74696d16da54 · outbound

This paper cites ” why should i trust you?” explaining the predictions of any classifier.

Token Activation Map to Visually Explain Multimodal LLMs ” why should i trust you?” explaining the predictions of any classifier

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.802753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:26.369433Z digest=sha256:8ec10eeaa50f9dd6570c1da7f52bebe9c6ad542195a67ebce4ec9e5bef37d36a

Observation 05a57b0f-e59c-410c-a9bb-cfb97897fd94 · outbound

This paper cites Causal interpretation of self-attention in pre-trained trans- formers.

Token Activation Map to Visually Explain Multimodal LLMs Causal interpretation of self-attention in pre-trained trans- formers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.626574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:26.613931Z digest=sha256:3678c7b4af6dbf83611abcf33d3bab8e7bf25542e886be18e84b511d2aaa8950

Observation 196618f3-66a4-4c43-aee8-be7d529c1fed · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

Token Activation Map to Visually Explain Multimodal LLMs Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:26.849975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:26.849975Z digest=sha256:ea06e17e50057438fc4962de5cffa203374c4170798b580aa077469d8ee34e68

Observation 67ed4cc3-904b-4d83-a56e-cd116cc1d3a2 · outbound

This paper cites Training- free object counting with prompts.

Token Activation Map to Visually Explain Multimodal LLMs Training- free object counting with prompts

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.256509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:27.179429Z digest=sha256:d7ea9298170f358e1fc15f0a52fb3f1eef58271fb6ef6659ac5a386046bd7254

Observation 1a7d0272-266c-4853-b2cc-4f7fc734a287 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Token Activation Map to Visually Explain Multimodal LLMs Gemini: A Family of Highly Capable Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:27.282270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:27.282270Z digest=sha256:f50878821ca474a37acb3e52faee74dfa987cdabc6b3b42a1941cc0d160db508

Observation 6150cd2f-2db5-48e3-bbb1-95b8b2d7738d · outbound

This paper cites Understanding how vision-language models rea- son when solving visual math problems.

Token Activation Map to Visually Explain Multimodal LLMs Understanding how vision-language models rea- son when solving visual math problems

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:32.082031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:27.416113Z digest=sha256:da74734900ef5d56e4bd75fff293c3c3ff73c55db3e83e24a3b5b6c9f5868bce

Observation 7a765326-b36b-454a-9900-2259ac8c7c62 · outbound

This paper cites Attention is all you need.

Token Activation Map to Visually Explain Multimodal LLMs Attention is all you need

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.920473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:27.564729Z digest=sha256:d2688f673f47925aabc7aaa8508873b6db924e27340375361c01871a4627ea10

Observation c69c4e1c-7ea3-477a-827b-b814833df753 · outbound

This paper cites Interpretable bilin- gual multimodal large language model for diverse biomed- ical tasks.

Token Activation Map to Visually Explain Multimodal LLMs Interpretable bilin- gual multimodal large language model for diverse biomed- ical tasks

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.726790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:27.740841Z digest=sha256:cd565a52dc144ca868aa76f10598a6dcde44797e07708cb67bb9e1aaddbac2b7

Observation f3bf8d0d-6f75-4ecd-baf9-b4b3a078a944 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Token Activation Map to Visually Explain Multimodal LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:27.901257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:27.901257Z digest=sha256:193e211021ddeb8771cb659981bea2a433cba4e33b7637f9aad81873406adcbc

Observation 1593e6e8-1fc5-4776-a1c7-7620dcf83565 · outbound

This paper cites Star: A benchmark for situated reasoning in real-world videos.

Token Activation Map to Visually Explain Multimodal LLMs Star: A benchmark for situated reasoning in real-world videos

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.519512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:28.031705Z digest=sha256:5ac88c37004105fec3ece5fc48dad448e3d917452c4eafee0749dab100926534

Observation 1e49da6f-40ab-42a7-aeb0-a8ae0eeb2f01 · outbound

This paper cites Efficient streaming language models with attention sinks.

Token Activation Map to Visually Explain Multimodal LLMs Efficient streaming language models with attention sinks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.306564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:28.184907Z digest=sha256:4f009d1acc7e90dfdb5bf3fea577f62b4b6f87858a9f80373ce1cb5ddcb0d23e

Observation 84be4017-1d5c-4212-b382-2eb7d37008b0 · outbound

This paper cites A survey on causal inference.

Token Activation Map to Visually Explain Multimodal LLMs A survey on causal inference

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:31.119616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:28.331544Z digest=sha256:98e676fda224d3c514701f54a281a0548bee0515466a7b984b0209f94ed4272f

Observation 0e52b13d-beb8-43b3-8e25-e2980e367942 · outbound

This paper cites From redundancy to relevance: Enhancing explainability in multimodal large language mod- els.

Token Activation Map to Visually Explain Multimodal LLMs From redundancy to relevance: Enhancing explainability in multimodal large language mod- els

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.950021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:28.453431Z digest=sha256:f9fbd220cf67db26b424dd771ba42c54a7537de9aad16d3915cb4c0ca3a370dc

Observation 0ec2ab28-29f7-4ab3-8a1a-d02842692370 · outbound

This paper cites Learning deep features for discrimina- tive localization.

Token Activation Map to Visually Explain Multimodal LLMs Learning deep features for discrimina- tive localization

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T21:52:28.534320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:52:28.534320Z digest=sha256:0a8aae025357e88a72fbd6148c176ddb7fd511c4d4fb57178e6bb25a25828965

Observation 9278089c-6546-41d2-b914-a25a9151eebb · outbound

This paper cites with” and the punctuation mark “.

Token Activation Map to Visually Explain Multimodal LLMs with” and the punctuation mark “

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.630735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:28.788264Z digest=sha256:e8c39f50e2692ab954dad2579da7cf50d003c05a237b453163de65e8474f5c40

Observation 6bf0e808-16c8-4475-b27e-9a9e1a53e8d0 · outbound

This paper cites an unresolved cited work.

Token Activation Map to Visually Explain Multimodal LLMs Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:30.447211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:28.941451Z digest=sha256:798457594b69acae30a8df56b378ec4b6fcd7336f1547c614bebeb22ef58cfb2

Observation 34a50edb-53d1-4e84-b268-dc2130cfdc34 · outbound

This paper cites living wall.

Token Activation Map to Visually Explain Multimodal LLMs living wall

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.265472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:29.101599Z digest=sha256:9bbad6c4f04c315cdd072f3a5d75ecc674b59b9cc4189d90b995d2d8e7570a77

Observation cad3c0e3-9e47-4c87-b66d-b1722867a254 · outbound

This paper cites living wall.

Token Activation Map to Visually Explain Multimodal LLMs living wall

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:30.081243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:29.230235Z digest=sha256:37f9766174886d47921861c24fbc0e993f59e885bc828a9d41d2979ef571820a

Observation 0ec197a3-1b0d-48b1-85b4-bda3adb45db0 · outbound

This paper cites Object-determined.

Token Activation Map to Visually Explain Multimodal LLMs Object-determined

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:29.923294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:29.318102Z digest=sha256:44ea7a18b66038914aaaca68539fa7a38904601c151edb5a1b7d832b43072746

Observation 42bfa086-e8c6-463a-8d82-7eb6d2113a8c · outbound

This paper cites Missing arrows led to erroneous reasoning.

Token Activation Map to Visually Explain Multimodal LLMs Missing arrows led to erroneous reasoning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:29.760939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:29.407583Z digest=sha256:81579dcb449fd5369e2d81d8eda9facb52d63a78c7b775f473bea0a448fd610c

Observation fa6da81a-05a1-4059-a00b-26a19f33bfe0 · outbound

This paper cites 2, 5, 6, 7, 14, 15.

Token Activation Map to Visually Explain Multimodal LLMs 2, 5, 6, 7, 14, 15

Reference 168

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:38.764893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:23.026379Z digest=sha256:f65dc2b061f1fcd4eefc667f82c45f27a34d73855c2d8f3114ee79e208e43989

Observation c8836885-99c1-429f-8b6c-71ada12ddd70 · outbound

This paper cites an unresolved cited work.

Token Activation Map to Visually Explain Multimodal LLMs Unresolved cited work

Reference 2016

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:30.757393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:28.664380Z digest=sha256:4750bdb981757530bfdbb03c6bd572096c4465e5fe33d0f2562026824960fe1c

Observation ea323d48-56d5-4c18-97cd-36edf6b02d54 · outbound

This paper cites an unresolved cited work.

Token Activation Map to Visually Explain Multimodal LLMs Unresolved cited work

Reference 2017

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:32.418675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:52:27.014153Z digest=sha256:e62f8ea7912b8ff59469ddebcbcc614fc3ffd1a6d71ce9c6c38d05ea8ad3ff82

Pith citing papers

Observation ee607a6c-424c-4103-8a9b-d5ac4adba379 · inbound

Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models cites this paper.

Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models Token Activation Map to Visually Explain Multimodal LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:38.554863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:58:38.554863Z digest=sha256:b84a9941f66a7f92aefc82b38672f4703655dd77f046beb0b5adbb404daa6cc0

Observation 61037376-a873-45fb-a438-086e6355c419 · inbound

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward cites this paper.

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward Token Activation Map to Visually Explain Multimodal LLMs

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:47.892061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:59:19.379119Z digest=sha256:e786cb1c74badc30f3bcb7bd90a4d6ea685f9dacfe82f5120472e9420d2c111a

Observation 68b4e094-aaed-4176-9567-dc7d2db5ea9d · inbound

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference cites this paper.

MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference Token Activation Map to Visually Explain Multimodal LLMs

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:45:52.037170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T01:29:26.298131Z digest=sha256:e04831b7ad344f9d108031d4fb6fa561e28e168d7cd14a094e62744cd566c4e2