Pith. sign in

Paper Citation Record · LEDGER

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering

As of 9 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 0 inbound Pith citation observations for arXiv:2507.12490.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12490 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:07:54.402782Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0f40e080-bd1d-4602-a9a9-299172ff8252 · outbound

This paper cites In: Proceedin gs of the IEEE/CVF international conference on computer vision.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Proceedin gs of the IEEE/CVF international conference on computer vision

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.878906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:07:52.422420Z digest=sha256:be8052eacc9324f69cd86d8db9f0277f1a4956ef536fd3a5c93ae702a1f43817

Observation 49e34a29-2d3c-479a-9316-6c6c8e621f93 · outbound

This paper cites 4290–4300.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering 4290–4300

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:52.513301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:52.513301Z digest=sha256:666be0b66e2c460719f7acd9aa6923ae9ca8421c3adb22b1426d215cfd55f80b

Observation 792ca61a-d574-44fc-bbd4-8aa07850e039 · outbound

This paper cites Pattern Recognition Letters 150, 242–249 (2021).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering Pattern Recognition Letters 150, 242–249 (2021)

Reference 3

Resolution
verified exact
doi, observed 2026-08-06T17:07:54.543867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:07:52.593231Z digest=sha256:447548fe20111065053eb7fa255a15d866525ecb0c177303173127089252887e

Observation 1e4266db-48e4-4e6e-90ed-a8f540d82797 · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:52.666510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:52.666510Z digest=sha256:b37625bcc1d57f53cbf5c317663113738d2103dc035375648b91adf13ecde28f

Observation 57f7fc36-5dae-4faa-aa87-e576c9420f82 · outbound

This paper cites In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recog- nition (CVPR)

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:52.747544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:52.747544Z digest=sha256:5506860b27c19d530e2be346301992d589ac25c860888236cbfbd0b0559c40ce

Observation 6b80b483-998d-4af9-8f1a-fd14e141652a · outbound

This paper cites , Le, Q., Sung, Y.H., Li, Z., Duerig, T.: Scaling up visual and vision-langu age representation learning with noisy text supervision.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering , Le, Q., Sung, Y.H., Li, Z., Duerig, T.: Scaling up visual and vision-langu age representation learning with noisy text supervision

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.649166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:07:52.849169Z digest=sha256:490c558bbfa76a7a63ebe603f3228644a0b9da55f30f665eee005b02f3f43bd8

Observation fe155ab4-0f54-40be-aba1-65d4068bdd88 · outbound

This paper cites In: European Conference on Computer Vision.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: European Conference on Computer Vision

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.509448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:07:52.991227Z digest=sha256:1ebe2001a95768844261f336859274d71c20df96876cd4d5cef6d1288503d215

Observation 2033bde0-f91e-484d-80fe-e51dd8abe144 · outbound

This paper cites In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlet t, J.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlet t, J

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.395421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:07:53.076926Z digest=sha256:832703083206b022597959b7859acc320b71694fa8b1f10bd8402687739631ab

Observation 63e3231a-10b2-406a-aa7c-06d68f5fc3de · outbound

This paper cites In: Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., Sabato, S.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., Sabato, S

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.258565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:07:53.150218Z digest=sha256:90c044941dc6ae611ac1473d376c9085307db65bc425f3d6ea0adafd73cd262c

Observation 64518d66-6287-431b-bb90-c0e23b7515cb · outbound

This paper cites Multimodal Rationales for Explainable Visual Question Answering.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering Multimodal Rationales for Explainable Visual Question Answering

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:07:54.956540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:07:53.224963Z digest=sha256:712411915575ef12a8b9a0d50a78a1bd02ef7e978cd9782435701b0fda3c4cbb

Observation 43fea144-a5dd-4d2c-8b09-ac79d5b0b30b · outbound

This paper cites In: Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:56.118590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:07:53.317817Z digest=sha256:da1c0d62582eaf48ce4db7afed21c4a0672fc4adef4c37d07d7c8cddee9e96a7

Observation 604c8031-6688-4ef7-9f19-cd35ea01d861 · outbound

This paper cites In: Proceedings of the IEEE/CVF winter conference o n applications of computer vision.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Proceedings of the IEEE/CVF winter conference o n applications of computer vision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.976919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:07:53.417551Z digest=sha256:4a204464b76671149e04e7b91ac7e90f01b4a7128102c61ee7e30a7f2497b6dc

Observation b8fcff72-8148-4bbc-8735-301968efb62e · outbound

This paper cites DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:53.544005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:53.544005Z digest=sha256:745a8aeb1eb6b7fd5aa45bca192bcf787701f6206d620de34651507f72e319fe

Observation f32de7c0-4fbb-423c-8390-52f65fd8c9d4 · outbound

This paper cites , Pietruszka, M., Pałka, G.: Going full-tilt boogie on document understanding with t ext-image-layout trans- former.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering , Pietruszka, M., Pałka, G.: Going full-tilt boogie on document understanding with t ext-image-layout trans- former

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.852534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:07:53.639104Z digest=sha256:85d384ec60e5857fc1e2cc1595b23a2910b7eab4c49bab1bd97be26651b38e1d

Observation 75f70c2c-573a-45c4-aa15-8ff741b25099 · outbound

This paper cites In: Meila , M., Zhang, T.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Meila , M., Zhang, T

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.576413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:07:53.830975Z digest=sha256:cdab0090d12c779afb0f2c5a55e131706c3fe3829214fdbcd629ced1425fa5a5

Observation f96db652-053a-411e-b546-c39eb88f723c · outbound

This paper cites an unresolved cited work.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:07:55.747430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:07:53.707881Z digest=sha256:1b75090df544748d36f9f07faa4c6eb35562f7691616ee3071c934c690d7b056

Observation 6b65e2db-c640-4bb8-8cb1-04bb8cdf930f · outbound

This paper cites In: 2019 IEEE/CVF Conference on Computer Vi sion and Pattern Recognition (CVPR).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2019 IEEE/CVF Conference on Computer Vi sion and Pattern Recognition (CVPR)

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:53.924438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:53.924438Z digest=sha256:c3699f65f09d16701521e7e46ff7c86f3042c21b332efc912c3f3ce638aaaa80

Observation e0075ca2-de6a-4548-8926-de07ac41acc3 · outbound

This paper cites I n: International Confer- ence on Document Analysis and Recognition.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering I n: International Confer- ence on Document Analysis and Recognition

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.448109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:07:54.023790Z digest=sha256:9678a093b593900b3017ffde1fec8fb9669a85b5e7ab78bc1e72ba979253460d

Observation 41c2d7c5-b59b-4419-b8d9-31c5b137352a · outbound

This paper cites In: 2017 IEEE International Conference on Computer Vision (ICC V).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2017 IEEE International Conference on Computer Vision (ICC V)

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:54.117543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:54.117543Z digest=sha256:a4d5ef592528857e6411e0538014d863de9930c643c0ea6264f80e0c573ef564

Observation 60e6df6a-3035-4fc7-ae12-9fe0bf21b657 · outbound

This paper cites In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recogni tion (CVPR).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recogni tion (CVPR)

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:54.246311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:54.246311Z digest=sha256:cdac4b0601e8026a6e7d8e8bb050dcea04f78dfbb6759da1bd1bf6a1797bdb0e

Observation 45eb9f88-24a5-4f21-9c0a-c4f896218a83 · outbound

This paper cites In: Proceedings of the 26th ACM SIGKDD International Confer - ence on Knowledge Discovery & Data Mining.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: Proceedings of the 26th ACM SIGKDD International Confer - ence on Knowledge Discovery & Data Mining

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:54.339787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:54.339787Z digest=sha256:5999014b4c76194b8e0951dc22a4c1148964af91bd36caf24716121d49b4afc9

Observation 1cf8540d-5b29-4241-8ae6-4ecb408f8a18 · outbound

This paper cites In: The E leventh International Conference on Learning Representations (2022).

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering In: The E leventh International Conference on Learning Representations (2022)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:07:55.243516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:07:54.402782Z digest=sha256:e1757000b2bd49f77b35dccaf42f939ffc6a99d460c89dd1d6b23bac3f18fc20

Pith citing papers

No inbound Pith citation observations are available.