Pith. sign in

Paper Citation Record · LEDGER

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering

As of 11 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2507.12490.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12490 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:07:54.402782Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0f40e080-bd1d-4602-a9a9-299172ff8252 · outbound

This paper cites In: Proceedin gs of the IEEE/CVF international conference on computer vision.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Proceedin gs of the IEEE/CVF international conference on computer vision

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.878906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T17:07:52.422420Z digest=sha256:3c5833be057d04f76c07b7a7e392647b88e785d8925c3f804a54efef1d7b83f0

Observation 49e34a29-2d3c-479a-9316-6c6c8e621f93 · outbound

This paper cites 4290–4300.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering 4290–4300

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:52.513301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:52.513301Z digest=sha256:666be0b66e2c460719f7acd9aa6923ae9ca8421c3adb22b1426d215cfd55f80b

Observation 792ca61a-d574-44fc-bbd4-8aa07850e039 · outbound

This paper cites Pattern Recognition Letters 150, 242–249 (2021).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering Pattern Recognition Letters 150, 242–249 (2021)

Reference 3

Resolution
verified exact
doi, observed 2026-08-06T17:07:54.543867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T17:07:52.593231Z digest=sha256:90101d425c38dae72acec47be2eeb9e49bcf75ca28127bbb7dc144727c9e3947

Observation 1e4266db-48e4-4e6e-90ed-a8f540d82797 · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:52.666510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:52.666510Z digest=sha256:536345d8969623d8d15eb8fceccccce3a89f6978af438c579116edbd2d641de5

Observation 57f7fc36-5dae-4faa-aa87-e576c9420f82 · outbound

This paper cites In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR)

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:52.747544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:52.747544Z digest=sha256:5506860b27c19d530e2be346301992d589ac25c860888236cbfbd0b0559c40ce

Observation 6b80b483-998d-4af9-8f1a-fd14e141652a · outbound

This paper cites , Le, Q., Sung, Y.H., Li, Z., Duerig, T.: Scaling up visual and vision-langu age representation learning with noisy text supervision.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering , Le, Q., Sung, Y.H., Li, Z., Duerig, T.: Scaling up visual and vision-langu age representation learning with noisy text supervision

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.649166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T17:07:52.849169Z digest=sha256:48428ed150c8544a8a3e0d96974374d01cad42b0512fdb5366efcd14a858559a

Observation fe155ab4-0f54-40be-aba1-65d4068bdd88 · outbound

This paper cites In: European Conference on Computer Vision.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: European Conference on Computer Vision

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.509448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T17:07:52.991227Z digest=sha256:828ce6668dc427780e802a7062e0c2cb484c9aac782dc5c72a837a63b4651f48

Observation 2033bde0-f91e-484d-80fe-e51dd8abe144 · outbound

This paper cites In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlet t, J.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlet t, J

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.395421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T17:07:53.076926Z digest=sha256:c3ff66f9d0d51732fbe267a5451d515e336b3324de196cd73a078e61171afe76

Observation 63e3231a-10b2-406a-aa7c-06d68f5fc3de · outbound

This paper cites In: Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., Sabato, S.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., Sabato, S

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.258565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T17:07:53.150218Z digest=sha256:2168c1bbde51b8b4ffacde116c0e4eb422a0b34cdffb761ece49c15bbaa7e1de

Observation 64518d66-6287-431b-bb90-c0e23b7515cb · outbound

This paper cites Multimodal Rationales for Explainable Visual Question Answering.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering Multimodal Rationales for Explainable Visual Question Answering

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:07:54.956540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T17:07:53.224963Z digest=sha256:60e70af032bd47c10b13b3330764d72942c8c35cce22e3f6edba94dca85a65ee

Observation 43fea144-a5dd-4d2c-8b09-ac79d5b0b30b · outbound

This paper cites In: Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.118590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T17:07:53.317817Z digest=sha256:abb15dcc426083a0338f1d74dffdeb769bfd10995aedbd8daf18ad20faa6786b

Observation 604c8031-6688-4ef7-9f19-cd35ea01d861 · outbound

This paper cites In: Proceedings of the IEEE/CVF winter conference o n applications of computer vision.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Proceedings of the IEEE/CVF winter conference o n applications of computer vision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.976919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T17:07:53.417551Z digest=sha256:38fbe42a3ffa50599ba5223e40465b5c0d7d26bbcdb020244663e644f7724c03

Observation b8fcff72-8148-4bbc-8735-301968efb62e · outbound

This paper cites DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:53.544005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:53.544005Z digest=sha256:debe5bc8975a48019c78a5f25a752ad749ac8ed5320a4fa5ace7382d47ab7c81

Observation f32de7c0-4fbb-423c-8390-52f65fd8c9d4 · outbound

This paper cites , Pietruszka, M., Pałka, G.: Going full-tilt boogie on document understanding with t ext-image-layout trans- former.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering , Pietruszka, M., Pałka, G.: Going full-tilt boogie on document understanding with t ext-image-layout trans- former

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.852534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T17:07:53.639104Z digest=sha256:9b8d1b94f39be9a8231f950f5e76e26ae03ce6f03f83af3a735c7cea30aa7ee3

Observation 75f70c2c-573a-45c4-aa15-8ff741b25099 · outbound

This paper cites In: Meila , M., Zhang, T.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Meila , M., Zhang, T

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.576413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T17:07:53.830975Z digest=sha256:98ac8ecb01809603cdf97ec6f92259f29072c16233d17add3e8b825d2f0a9fe9

Observation f96db652-053a-411e-b546-c39eb88f723c · outbound

This paper cites an unresolved cited work.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:07:55.747430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T17:07:53.707881Z digest=sha256:92ad0e31e0c78d73faeaa69e02589615df47575a3755a2be547c2c4827d78896

Observation 6b65e2db-c640-4bb8-8cb1-04bb8cdf930f · outbound

This paper cites In: 2019 IEEE/CVF Conference on Computer Vi sion and Pattern Recognition (CVPR).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2019 IEEE/CVF Conference on Computer Vi sion and Pattern Recognition (CVPR)

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:53.924438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:53.924438Z digest=sha256:c3699f65f09d16701521e7e46ff7c86f3042c21b332efc912c3f3ce638aaaa80

Observation e0075ca2-de6a-4548-8926-de07ac41acc3 · outbound

This paper cites I n: International Confer- ence on Document Analysis and Recognition.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering I n: International Confer- ence on Document Analysis and Recognition

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.448109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T17:07:54.023790Z digest=sha256:e8c31d9ae5da53b0ad06a999d19831e0d48307586e69c82ac447d5743173d32c

Observation 41c2d7c5-b59b-4419-b8d9-31c5b137352a · outbound

This paper cites In: 2017 IEEE International Conference on Computer Vision (ICC V).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2017 IEEE International Conference on Computer Vision (ICC V)

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:54.117543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:54.117543Z digest=sha256:a4d5ef592528857e6411e0538014d863de9930c643c0ea6264f80e0c573ef564

Observation 60e6df6a-3035-4fc7-ae12-9fe0bf21b657 · outbound

This paper cites In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recogni tion (CVPR).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recogni tion (CVPR)

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:54.246311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:54.246311Z digest=sha256:cdac4b0601e8026a6e7d8e8bb050dcea04f78dfbb6759da1bd1bf6a1797bdb0e

Observation 45eb9f88-24a5-4f21-9c0a-c4f896218a83 · outbound

This paper cites In: Proceedings of the 26th ACM SIGKDD International Confer - ence on Knowledge Discovery & Data Mining.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Proceedings of the 26th ACM SIGKDD International Confer - ence on Knowledge Discovery & Data Mining

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:54.339787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:54.339787Z digest=sha256:5999014b4c76194b8e0951dc22a4c1148964af91bd36caf24716121d49b4afc9

Observation 1cf8540d-5b29-4241-8ae6-4ecb408f8a18 · outbound

This paper cites In: The E leventh International Conference on Learning Representations (2022).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: The E leventh International Conference on Learning Representations (2022)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.243516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-06T17:07:54.402782Z digest=sha256:30a90966ee5fa3b61e58948296fa3fb2ced3fc1e4bfe3737fcae18aea92b95b9

Pith citing papers

No inbound Pith citation observations are available.