Pith. sign in

Paper Citation Record · LEDGER

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models

As of 21 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 2 inbound Pith citation observations for arXiv:2506.18985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18985 v3

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:46:03.748195Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T19:59:19.379119Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact2
  • verified fuzzy15
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 9368b046-2675-46dd-b26f-c9dd94caf847 · outbound

This paper cites Quantifying attention flow in transformers.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Quantifying attention flow in transformers

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.117975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.639483Z digest=sha256:e931f3acb8f5513a762812b373e30b6f111be625f738e79d3f94528d7e979846

Observation dad3ad63-80a2-480c-a477-43c140ab9e35 · outbound

This paper cites AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:03.643989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:03.643989Z digest=sha256:2bd6639279e09e03189de8f19d64d87eaa0999c7164f04317f81c3c4e2d5e12f

Observation bfce766f-b504-4d1e-ac24-49f90eba467d · outbound

This paper cites XAI for trans- formers: better explanations through conservative propaga- tion.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models XAI for trans- formers: better explanations through conservative propaga- tion

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.105967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.648241Z digest=sha256:c01709c802ba94daea9c197b3d8be593ebeab12923d766bd5597d8963233dfcb

Observation c302c794-ca24-4bc0-8f2d-7ee72c1b7db3 · outbound

This paper cites On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.093686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.652167Z digest=sha256:61e202a7a6c290f21f4e798aef3a456bb385e4e736ba7d9a0fbcb646844a3391

Observation cc955556-c5c7-42e0-b57a-da82e4a71a53 · outbound

This paper cites Qwen2.5-VL Technical Report.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Qwen2.5-VL Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:03.656596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:03.656596Z digest=sha256:b7cc6ce0a8734980de1b72c098da0b9439db92a3774d47a2c6742524a0a4a374

Observation f17b3e68-7802-453c-93ee-23e3b4df9b09 · outbound

This paper cites an unresolved cited work.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:46:04.082390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.660515Z digest=sha256:2da3332fe4816b519504db59067df22cc4ecbcfd1a8be248bddd1adc9ae71ca1

Observation b3a63cf5-3d53-471f-bc0b-0ce2f469bbae · outbound

This paper cites Visual Explanations via Iterated Integrated Attributions.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Visual Explanations via Iterated Integrated Attributions

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:46:03.908685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.664989Z digest=sha256:1515fb972509f59d69cc0b40cd55555ef852f95ee33b463171d2cf54ca319dd6

Observation 1eedf591-3e3b-409f-b668-3ac954c14642 · outbound

This paper cites Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.071386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.669555Z digest=sha256:cc52cffec70fec5f7b4b88957558543e25c793055a04877813388ad72313086f

Observation 8a60b89c-3f04-4ab7-bed2-be1f59881448 · outbound

This paper cites Human attention in visual question answer- ing: do humans and deep networks look at the same regions? In Proc.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Human attention in visual question answer- ing: do humans and deep networks look at the same regions? In Proc

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.060196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.673042Z digest=sha256:f57dc50bccfbe0c5577d4f97bd5a907aaee92f7561fe5f061ded75d2e165b306

Observation 3698a970-62c3-4bf5-bbd6-1a3a5a009400 · outbound

This paper cites AtMan: Understanding Transformer Predictions Through Memory Efficient Attention Manipulation.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models AtMan: Understanding Transformer Predictions Through Memory Efficient Attention Manipulation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:03.676556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:03.676556Z digest=sha256:fb0aa0fb176c07000f3778b3cb736035f1dce3f8336dec91ce804c90943cf6cc

Observation 3b3c9461-e966-43be-ac55-28178f7d9727 · outbound

This paper cites Turtles, Hats and Spectres: Aperiodic structures on a Rhombic tiling.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Turtles, Hats and Spectres: Aperiodic structures on a Rhombic tiling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:03.680944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:03.680944Z digest=sha256:fd8f974b096a8a715427e32fe4c6b1bfb37378938b1985cbeadf3f97b1610f9f

Observation 49838183-700a-4569-8e91-10c49ea97e7d · outbound

This paper cites iGOS++: inte- grated gradient optimized saliency by bilateral perturbations.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models iGOS++: inte- grated gradient optimized saliency by bilateral perturbations

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.048896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.684905Z digest=sha256:eda0fad03e82a03924993c9fa4d3309a463052571e10e1a106e59166cab4b151

Observation 4948d703-030d-414c-8f24-d2cee8ca793c · outbound

This paper cites From representa- tion to reasoning: Towards both evidence and commonsense reasoning for video question-answering.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models From representa- tion to reasoning: Towards both evidence and commonsense reasoning for video question-answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.037962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.688639Z digest=sha256:88e76f2c26c778ebc057173f4a71bb4a5a85f884719d05055fec0fd10c51e9de

Observation 0925e4f2-94ce-49c2-a15e-37135eb8dbfb · outbound

This paper cites Visual Instruction Tuning.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Visual Instruction Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:03.692170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:03.692170Z digest=sha256:b360ecc0a7cf73517a11a3022f3e92d5d97f10cbc6191d649caeb3adc3303fd3

Observation b30bf151-85f1-44f2-9379-fe6a621791d2 · outbound

This paper cites Lundberg and Su-In Lee.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Lundberg and Su-In Lee

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.026785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.696219Z digest=sha256:2d2c323c4f2e5ce306e367c65749ecbaaa4eab13d5a68491eebc49809097f2c6

Observation 6f305ed1-19a6-4792-aff2-1ffc1df838cc · outbound

This paper cites Mortality Forecasting using Variational Inference.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Mortality Forecasting using Variational Inference

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T18:46:03.860984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.700734Z digest=sha256:60dd763ed38291ee8e8c075e01e5bb9df2c5fc2be926eb92172d496c299aca6b

Observation 8625d7cc-78aa-4e18-9c69-5d185cb40c03 · outbound

This paper cites Exploring human-like attention supervision in visual question answer- ing.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Exploring human-like attention supervision in visual question answer- ing

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.015322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.705261Z digest=sha256:59c0aea9c62668ffc7645eb1f7a45d8af9c4b680f1a0dff6a7297585d39acb22

Observation 8e4f297d-0d47-4626-af6e-e971c00aff3c · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:03.708856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:03.708856Z digest=sha256:2270b3f4b4831aac285086acf3820dcce417208d86adf95a7ebb7d56d989e666

Observation 85c84dc7-508d-4685-8e72-1e1ea0405777 · outbound

This paper cites Uncertainty estimates for semantic segmentation: providing enhanced reliability for automated motor claims handling.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Uncertainty estimates for semantic segmentation: providing enhanced reliability for automated motor claims handling

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T18:46:03.833662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.712921Z digest=sha256:5b98d7636dee8ddb0d8084007f60d226e54c08cc1d23c9509c9233dc3f02908e

Observation 4810cd64-9b2c-47a6-a56f-ca0ebda08dab · outbound

This paper cites Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Ba- tra.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Ba- tra

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:04.002676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.716980Z digest=sha256:1b92704dcd59ee492ad0e6a93f7e4f13cc5451dfeb550d83f2b6973802bc9afd

Observation ab0d3a7d-9336-465e-83f1-0972c0f63b06 · outbound

This paper cites Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Ba- tra.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Ba- tra

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:03.989684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.722298Z digest=sha256:b82ab5b52a6a2cd36e7be28edcad72ceeb75e088ed01e6c7568f6aef725e7ab1

Observation 13db61c4-7aed-41f1-821b-af02a44434df · outbound

This paper cites Deep inside convolutional networks: visualising image clas- sification models and saliency maps.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Deep inside convolutional networks: visualising image clas- sification models and saliency maps

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:03.976671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.725942Z digest=sha256:b6a3000d9a2be2bd5f40ce3b6cf5c9e36dc363bcaff2125770c6b8601705ffd1

Observation d3968f7e-926d-4681-8dc7-aa014e420dcc · outbound

This paper cites On numerical solutions of the time-dependent Schr\"odinger equation.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models On numerical solutions of the time-dependent Schr\"odinger equation

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T18:46:03.814922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.729572Z digest=sha256:203a356bf6b7084be748e57b2ecee1c8b73d9ea28f8e13948c8b9c3b1a767ff4

Observation 01581318-b063-4710-9727-44e317e01030 · outbound

This paper cites LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T18:46:03.733512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:46:03.733512Z digest=sha256:e1c3086526b14b7386b7b2369b58ba9142cf253b59e976a2f7e8f5ca78497cb6

Observation ef6f40f4-bdf6-47ab-8fde-1911bff2fc9b · outbound

This paper cites Axiomatic attribution for deep networks.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Axiomatic attribution for deep networks

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:03.964230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.737200Z digest=sha256:f66280092b935bcd2cf43f4c0e4089cec0f3b7dcf19e157684f2a7b9893b9fd7

Observation 5c4d78c7-4ea7-4683-866c-3774b2dd847d · outbound

This paper cites Attention, Please! PixelSHAP Reveals What Vision-Language Models Actually Focus On.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models Attention, Please! PixelSHAP Reveals What Vision-Language Models Actually Focus On

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:46:03.785591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.740939Z digest=sha256:82a3e3905ef438b0f3ee13d9ac15955d4c0324609f6d7e316a925293c4a587c4

Observation 9758877d-3a24-40fc-b4b4-0f5d28450645 · outbound

This paper cites VQA-MHUG: human gaze supervision for visual question answering.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models VQA-MHUG: human gaze supervision for visual question answering

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:03.952722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.744697Z digest=sha256:32f90be3c1a4c922ecad572c353c0998ea17bec30ff9f4cbd62b5497d8729222

Observation b27832b6-71d2-45dd-9a4b-b6ff578538dc · outbound

This paper cites What if the tv was off? examining counterfactual reasoning abilities of vision-language models.

GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models What if the tv was off? examining counterfactual reasoning abilities of vision-language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:46:03.940691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T18:46:03.748195Z digest=sha256:1531a0c1ad8272e26386c9e13de6f953b1fdabb2f81eb7f1f7d201c0b9b3b251

Pith citing papers

Observation 2bfed174-41f7-409f-974f-a06262405880 · inbound

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward cites this paper.

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:47.749240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T19:59:19.379119Z digest=sha256:57e3d6b39d2bfd0240c517fb58753c1243858c2da62336dff456828fcbe1a329

Observation 79adc347-769b-4532-bf89-f49ad36076e7 · inbound

Through Their Eyes: Fixation-aligned Tuning for Personalized User Emulation cites this paper.

Through Their Eyes: Fixation-aligned Tuning for Personalized User Emulation GLIMPSE: Holistic Cross-Modal Explainability for Large Vision-Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:35:47.332353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:21:47.577839Z digest=sha256:8953a66156fea0b7d95413c608ae9e3b94493cac6b2c532a6553f9a7fb3c2bdd