Pith. sign in

Paper Citation Record · LEDGER

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness

As of 13 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 3 inbound Pith citation observations for arXiv:2412.00151.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00151 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:13:29.887857Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:07:53.544005Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T04:44:02.457882Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e0d80dc3-aaf1-4e20-a0dd-f274cd6183cb · outbound

This paper cites write newline.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.686393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.686393Z digest=sha256:c5b54435806afdec39dab76eb737ce0d07449a822bf2931003cd44a7f271d37a

Observation 9bcb9371-e82c-4bc1-8af1-7e0b7714cf15 · outbound

This paper cites Pixtral 12B.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Pixtral 12B

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.764154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.764154Z digest=sha256:6a02cb6f5fd9147035a54beb3ba8169d7f17a2fa712e02f1487abfd6068b1853

Observation 0f71ecc7-3cfa-419d-b13c-962ce18621af · outbound

This paper cites Vision transformer for fast and efficient scene text recognition.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Vision transformer for fast and efficient scene text recognition

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:31.546888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:13:28.779678Z digest=sha256:2fcc4be4ea6cf48441d3baf61c8a51a87ea27f32503fe3bb804e97d43e0e1b5d

Observation 233e36ff-921b-4862-861c-7d5b6cd666e7 · outbound

This paper cites Scene text recognition with permuted autoregressive sequence models.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Scene text recognition with permuted autoregressive sequence models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:31.489758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:13:28.791634Z digest=sha256:6d34a6459374e8d4cdb746c4119ff8a4cc1969a681fd455d409769d2a7291e79

Observation fdf3778b-c7ff-4cc8-9bd0-9c1251f9c726 · outbound

This paper cites FAST: Faster Arbitrarily-Shaped Text Detector with Minimalist Kernel Representation.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness FAST: Faster Arbitrarily-Shaped Text Detector with Minimalist Kernel Representation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.799767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.799767Z digest=sha256:1b4cbfaab953f7af90e78159576cd5bf9b64b502b681927a016214b3c5dd76b1

Observation 8e4be6ac-9f30-4f9f-be78-d302dcc71de1 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.811476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.811476Z digest=sha256:548344ebf1955a9cf04cf3d59077dc27d71e9a557f86aca4fa567b7b2548e284

Observation 4444bbda-73bc-4e82-872f-3be4203913e7 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.841407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.841407Z digest=sha256:57b3ec511c39263f0ed54a5b889bc352c378b67f2be4f671ea2b43e984a02dc7

Observation 83f3f23e-022b-4587-8287-c376b7ab17cd · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.859207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.859207Z digest=sha256:9d80b45e577ba3296b1298b058f91efd3408fe3c1937de2625d04bc671a081d4

Observation 709d6934-67e5-454f-b908-494e30aa8bd3 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Qlora: Efficient finetuning of quantized llms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.868037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.868037Z digest=sha256:51f561b23c628f70ed9ee07e4d36c421b8c880d3bcb6bb7fb6ab0f6bddd26606

Observation ece1e68a-f183-4f67-af7a-b307804581c2 · outbound

This paper cites The Llama 3 Herd of Models.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.873924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.873924Z digest=sha256:3ac82e4ff7dbb8d6f0a0304b2739d04c5403362de436862f97e6b1c7ea3a9c95

Observation a6ff43f4-b7f9-4c69-8ed1-f37a2a170e15 · outbound

This paper cites Dtrocr: Decoder-only transformer for optical character recognition.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Dtrocr: Decoder-only transformer for optical character recognition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:31.400336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:13:28.879841Z digest=sha256:5b374449f552eeec3b820c74d7e333eeaf56cd168f843b0cb00cd94c785a8287

Observation f1f9abc0-e275-4b8e-a3e3-aae723b5c98d · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.886064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.886064Z digest=sha256:f5c4682b51e4f41d6d429dffc05004f72800c5f2c7053175a10637dcce969b52

Observation e7da7fed-e7a0-485d-b614-a55109a7b570 · outbound

This paper cites LoRA+: Efficient Low Rank Adaptation of Large Models.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness LoRA+: Efficient Low Rank Adaptation of Large Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.891759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.891759Z digest=sha256:3fdf1be9a2e71edf38638362254a705ee9aaf6d2630200ba23672afd24832bae

Observation 3cea5e5f-0040-4241-a0cd-613349378ddc · outbound

This paper cites Icl-d3ie: In-context learning with diverse demonstrations updating for document information extraction.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Icl-d3ie: In-context learning with diverse demonstrations updating for document information extraction

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:31.282162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:13:28.896935Z digest=sha256:4cb24c2ead78144ab8347f8a35e0f942e4245651ad5887ef5b1e994940071c9d

Observation 46c45ae0-8534-46bd-8d5c-3983522dec5a · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness LoRA: Low-Rank Adaptation of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.921379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.921379Z digest=sha256:c447f9c3461f950ea0a25c9f276de87515084067bab1e73069e996794f0b6b7b

Observation fd3eeaa5-026f-4530-98cc-0a5118ec88b8 · outbound

This paper cites Layoutlmv3: Pre-training for document ai with unified text and image masking.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Layoutlmv3: Pre-training for document ai with unified text and image masking

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:31.264667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.037690Z digest=sha256:c81cb900f761865e995cc3cf194056c3c15bc0bc80b6393e6f086ab1486b1f6b

Observation 464c226c-995c-479a-94e3-24e3fd79c8cc · outbound

This paper cites TrustLLM: Trustworthiness in Large Language Models.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness TrustLLM: Trustworthiness in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.102929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.102929Z digest=sha256:c6e6f8d873e0cb6d80d364a65864a4ae9265c84a99d70f9c1b83c9226904a4c0

Observation 4fc0f8b1-a59b-4dff-91f2-a8596553d501 · outbound

This paper cites Icdar2019 competition on scanned receipt ocr and information extraction.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Icdar2019 competition on scanned receipt ocr and information extraction

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:31.156302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.195342Z digest=sha256:ddd33de6222cbc7b22bbe5775a9bb6b96126b3c28d41c0a8426b5b4677a4bdca

Observation 33ce8a36-0753-4860-9bcc-56f1e33e89bd · outbound

This paper cites From image to language: A critical analysis of visual question answering (vqa) approaches, challenges, and opportunities.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness From image to language: A critical analysis of visual question answering (vqa) approaches, challenges, and opportunities

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:31.094217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.220974Z digest=sha256:8f95c47fef0dd07479e6785c97e8daa55d957a0a4f4caca371e04c47c5347ef6

Observation c289804d-c757-4bc5-8362-939f304538a0 · outbound

This paper cites Funsd: A dataset for form understanding in noisy scanned documents.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Funsd: A dataset for form understanding in noisy scanned documents

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:31.077837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.240067Z digest=sha256:2b8e43062eaff2ab8f66d369a14bacc8b270b4304722e319a729fce5f60e5f36

Observation c902eb2d-3712-4613-b96b-5fc156ac7031 · outbound

This paper cites Ocr-free document understanding transformer.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Ocr-free document understanding transformer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:31.061211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.260382Z digest=sha256:883e1d7094bafcffb20c3d11f7924627307a32e3b2a5374504e72dc70b699b31

Observation 095cf3f5-8ab4-44e6-9364-1bffa057965d · outbound

This paper cites Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language Models.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.276604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.276604Z digest=sha256:eff5059adbe5c5f553310a76e89b39ae71e2ccb71470ebeccf5716dd7f3fc181

Observation 91cd93d8-0d02-44e7-96ff-8281be9747c7 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness LLaVA-OneVision: Easy Visual Task Transfer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.281806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.281806Z digest=sha256:f1736792e7acd0918b7d93e93c007b6608b041f7252a9c9c3dcc076a070b66a6

Observation 6bcf163b-ccfa-429d-840b-cf3c19746ea1 · outbound

This paper cites Show, attend and read: A simple and strong baseline for irregular text recognition.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Show, attend and read: A simple and strong baseline for irregular text recognition

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.906380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.287419Z digest=sha256:d67b453987fe708598a16666adcc9e1a72f411c932e358e43f7d7a2876417850

Observation 14374c31-e293-4bdf-9370-5253ac943455 · outbound

This paper cites Trocr: Transformer-based optical character recognition with pre-trained models.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Trocr: Transformer-based optical character recognition with pre-trained models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.869327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.292071Z digest=sha256:e1e476bda563c17da29bb36694753d86f528fec4f916f20d1259367d281930c4

Observation 0a23538a-cf0a-4bb4-8c83-487443b8c7a2 · outbound

This paper cites Real-time scene text detection with differentiable binarization.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Real-time scene text detection with differentiable binarization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.854445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.297851Z digest=sha256:ddb363210ceea033bcd43f72dc2968a3328082a500bb3f4c8888921ff3a407d3

Observation cea610ae-b9d6-4768-9db0-8826fbc32897 · outbound

This paper cites DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.303009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.303009Z digest=sha256:cf3a82b46be3d2d3e9724097a1d556fab100cba9d809e8d1c013200dfdb160cb

Observation 04012c68-cb70-46e2-bc58-f28a534908ac · outbound

This paper cites DoRA: Weight-Decomposed Low-Rank Adaptation.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness DoRA: Weight-Decomposed Low-Rank Adaptation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.308712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.308712Z digest=sha256:fd16ae96539295ba6ddcbba6ef94aa65308edd4bd3f6f273361014cde07667ef

Observation 24918f95-f5ac-4e13-a9b6-509d24b8efd8 · outbound

This paper cites A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.314172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.314172Z digest=sha256:31a5a7e7b42f371ad05b138ab1828569411c2af490cd4c6be9100bb3cce874f8

Observation 41adb064-175b-4bf1-9bd7-8b884ddd6f79 · outbound

This paper cites Master: Multi-aspect non-local network for scene text recognition.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Master: Multi-aspect non-local network for scene text recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.832727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.319771Z digest=sha256:cd0af5e81389b32e6173e32c381eb3c49dcac9d4a9f8df2320cbbe7752a95cbe

Observation 3c7031a7-8f4a-41f9-9247-c52d6a116b24 · outbound

This paper cites Layoutllm: Layout instruction tuning with large language models for document understanding.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Layoutllm: Layout instruction tuning with large language models for document understanding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.666353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.324860Z digest=sha256:8e3f0e2060531b1430a0f9133fe3a43981142971517967e4469557875ef96856

Observation c55309fc-07ea-4dcf-ae19-6aef6a5878c7 · outbound

This paper cites MaskOCR: Text Recognition with Masked Encoder-Decoder Pretraining.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness MaskOCR: Text Recognition with Masked Encoder-Decoder Pretraining

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.397549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.397549Z digest=sha256:be24aaa9177992863fe21208c615b67cb965a091f515dbc33dfc1d0d7eb13850

Observation e8a5b922-66ea-488d-a05f-58ba5331d102 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Docvqa: A dataset for vqa on document images

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.484467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.484467Z digest=sha256:6ebcfe1c74dc16cfff0fb80732cca6fe1b29b2f5f54e5a3268fd092c9aca6816

Observation db6c78e2-d984-4198-a42c-2a2ef5d21ec1 · outbound

This paper cites Cord: a consolidated receipt dataset for post-ocr parsing.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Cord: a consolidated receipt dataset for post-ocr parsing

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.620032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.620032Z digest=sha256:b4f0e9ac404e079673b1b14f54c7f66d9e1178b1e5bd0f2eaf3e44a101fbb8d5

Observation 046668fb-9f94-48ed-900a-fea9b353c07c · outbound

This paper cites Generalized intersection over union: A metric and a loss for bounding box regression.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Generalized intersection over union: A metric and a loss for bounding box regression

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.486734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.711859Z digest=sha256:3bbaaf3a4b50630cb0997cbee1937a40623f02dd9bf23fb30268c785d5143c99

Observation 0dff8af7-c726-4a5c-b031-9daa5ef25106 · outbound

This paper cites An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.461824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.764976Z digest=sha256:fa03a194dc8a4902b1cb0dd0f446b27727607635d106983979dcd7f48443fd9c

Observation 09931be9-998f-4adf-8e53-60553123597a · outbound

This paper cites Instructdoc: A dataset for zero-shot generalization of visual document understanding with instructions.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Instructdoc: A dataset for zero-shot generalization of visual document understanding with instructions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.437769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.792352Z digest=sha256:a91220fdec4f84ea047efbd50e3251ff54dbbf67dc138430fc2451a538cec5fd

Observation 67b0da32-b059-4f5f-a8f7-adba54f3ba71 · outbound

This paper cites Unifying vision, text, and layout for universal document processing.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Unifying vision, text, and layout for universal document processing

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.390476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.825256Z digest=sha256:b18b205046eb882393f0d07f2efb6cf912ac8ed16d239cee213e0d65efbb2c7f

Observation 0138eb19-e9a8-4353-8477-4af778f08be8 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.848452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.848452Z digest=sha256:3302a5d19e5364f84392afb79ac8e6d840440c64bd8ad62c16827210667a2b2f

Observation ec38c8da-6808-4311-8cf9-b1aed03653b8 · outbound

This paper cites Omniparser: A unified framework for text spotting key information extraction and table recognition.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Omniparser: A unified framework for text spotting key information extraction and table recognition

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.348538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.858078Z digest=sha256:bf9213a381dfe9427cca587d72040f4fe8e8a70d8e396c1867e7c3453efc736d

Observation 56d10951-a55a-4939-ab07-6aa3e32019f4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.869788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.869788Z digest=sha256:f6ade1c53cb99563e0bbac97f4fb2f8126290294edfaca36b1edcc23c3ea6212

Observation ba2aeb0b-fcf8-4247-bb6e-ffd64ea540cf · outbound

This paper cites Layout and Task Aware Instruction Prompt for Zero-shot Document Image Question Answering.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Layout and Task Aware Instruction Prompt for Zero-shot Document Image Question Answering

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.874227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.874227Z digest=sha256:3cf8143124444405a41d2f7d5a4b4425438edaa8d748fb1b45c59a2915035ec4

Observation 938f48a7-1a9f-4219-b363-e08acc3aa3d5 · outbound

This paper cites A normalized levenshtein distance metric.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness A normalized levenshtein distance metric

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.878827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.878827Z digest=sha256:b811ff8c75087c15617bf138af709b3e950900dd0d7b66fb1c419e61cf1e5666

Observation d140bdda-db4f-4397-b489-6a1a58f4cafe · outbound

This paper cites MixNet: Toward Accurate Detection of Challenging Scene Text in the Wild.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness MixNet: Toward Accurate Detection of Challenging Scene Text in the Wild

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.882984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.882984Z digest=sha256:a6076a0dbec77f796390f077039d6306fa0eba9e5b6e848ebdc7d1a75edc71d2

Observation 7edc6aaf-2860-41f6-8347-58a8f0931ca9 · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.887857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.887857Z digest=sha256:34ae4aba969e34af8571fb84bf2722ca22f3b6895d0b2ad0b3948d3c72048ae9

Pith citing papers

Observation b8fcff72-8148-4bbc-8735-301968efb62e · inbound

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering cites this paper.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:53.544005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:53.544005Z digest=sha256:92ad9582578ca319b4c1faf93bc633ac8d8498d798144f45d56b6f6f091cc37f

Observation 8e5826c9-dd94-46cc-8d8c-4fe1a8678444 · inbound

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA cites this paper.

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:44:02.459977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-17T04:41:55.030673Z digest=sha256:014e6cc23c7898add008c88815dac951a174c7a7fad0afbe297ffc12e9765757

Observation 13c8431a-09e6-43fe-bf40-f1ecb82f34a6 · inbound

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment cites this paper.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:28.323502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:28.323502Z digest=sha256:16aa5ef4285c33a95b96ab09ee70909f8e980f927793b36d474df672bfcc2ab7