Pith. sign in

Paper Citation Record · LEDGER

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness

As of 14 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 3 inbound Pith citation observations for arXiv:2412.00151.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.00151 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:13:29.887857Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:07:53.544005Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T04:44:02.457882Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e0d80dc3-aaf1-4e20-a0dd-f274cd6183cb · outbound

This paper cites write newline.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.686393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.686393Z digest=sha256:b971e6c78ee7f6229841f9b995b39d2403524dcf5c39ee0e39a6b9151e3f7a9e

Observation 9bcb9371-e82c-4bc1-8af1-7e0b7714cf15 · outbound

This paper cites Pixtral 12B.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Pixtral 12B

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.764154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.764154Z digest=sha256:e829cc038904e91870113672689b08b74cb0c1e62ccb7dde9d30f73a1e8ef2f4

Observation 0f71ecc7-3cfa-419d-b13c-962ce18621af · outbound

This paper cites Vision transformer for fast and efficient scene text recognition.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Vision transformer for fast and efficient scene text recognition

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:31.546888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T10:13:28.779678Z digest=sha256:b69959b7ff96cb9e48e7dcf55d2245c26f1b682000e283eabe1f716b0a121cf7

Observation 233e36ff-921b-4862-861c-7d5b6cd666e7 · outbound

This paper cites Scene text recognition with permuted autoregressive sequence models.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Scene text recognition with permuted autoregressive sequence models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:31.489758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T10:13:28.791634Z digest=sha256:229d523bc1129b6671d5461c688d4eaa74587b202f67d4ab4bd626f676d134e3

Observation fdf3778b-c7ff-4cc8-9bd0-9c1251f9c726 · outbound

This paper cites FAST: Faster Arbitrarily-Shaped Text Detector with Minimalist Kernel Representation.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness FAST: Faster Arbitrarily-Shaped Text Detector with Minimalist Kernel Representation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.799767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.799767Z digest=sha256:b6cbd5392cd4b5d135f33b15bef72273c1ed4f4f33ed1511be806251249eecd0

Observation 8e4be6ac-9f30-4f9f-be78-d302dcc71de1 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.811476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.811476Z digest=sha256:a7d21646d36677c90ad3913acddf7010b79364e2f13a147e756b9636122c20ac

Observation 4444bbda-73bc-4e82-872f-3be4203913e7 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.841407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.841407Z digest=sha256:af7d472f3a644c715b740f9b83be19ce77606fa28930a1b6596712e1bdbbfc5f

Observation 83f3f23e-022b-4587-8287-c376b7ab17cd · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.859207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.859207Z digest=sha256:d24bb9d1d052a7dd8a41d45fd0bca379835812b5ae32cea003852f9aa632024f

Observation 709d6934-67e5-454f-b908-494e30aa8bd3 · outbound

This paper cites Qlora: Efficient finetuning of quantized llms.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Qlora: Efficient finetuning of quantized llms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.868037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.868037Z digest=sha256:324c2785dc598603452c7bfb4ed6c9d682237a4c73edbabc772e0551fc8a8d1f

Observation ece1e68a-f183-4f67-af7a-b307804581c2 · outbound

This paper cites The Llama 3 Herd of Models.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.873924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.873924Z digest=sha256:d561a8ae683a2d4026a28aa667739e8ae5e780ef414df3eef92452aaa502df5b

Observation a6ff43f4-b7f9-4c69-8ed1-f37a2a170e15 · outbound

This paper cites Dtrocr: Decoder-only transformer for optical character recognition.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Dtrocr: Decoder-only transformer for optical character recognition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:31.400336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T10:13:28.879841Z digest=sha256:73fe65c2afc19dfc4757ea32b4dbf2bb10132ea69a87c78b70486c44111b6b06

Observation f1f9abc0-e275-4b8e-a3e3-aae723b5c98d · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.886064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.886064Z digest=sha256:5773e4a8ee18f81ec944d314180f12d644da2862f2fc91fc1cc0c87ad3397be8

Observation e7da7fed-e7a0-485d-b614-a55109a7b570 · outbound

This paper cites LoRA+: Efficient Low Rank Adaptation of Large Models.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness LoRA+: Efficient Low Rank Adaptation of Large Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.891759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.891759Z digest=sha256:de86cb6a4678dfea519674c6562f447250d3155c06d490f5481d8ad125bf74a5

Observation 3cea5e5f-0040-4241-a0cd-613349378ddc · outbound

This paper cites Icl-d3ie: In-context learning with diverse demonstrations updating for document information extraction.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Icl-d3ie: In-context learning with diverse demonstrations updating for document information extraction

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:31.282162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T10:13:28.896935Z digest=sha256:d5a311e9462d9bdeba8da78f8f4b79a38a19a9b63b492fb0f0820425e819c978

Observation 46c45ae0-8534-46bd-8d5c-3983522dec5a · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness LoRA: Low-Rank Adaptation of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:28.921379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:28.921379Z digest=sha256:9fd7e39c3a8fd06746933d685399e670eb7941e933bda98cfc5caaa491105b81

Observation fd3eeaa5-026f-4530-98cc-0a5118ec88b8 · outbound

This paper cites Layoutlmv3: Pre-training for document ai with unified text and image masking.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Layoutlmv3: Pre-training for document ai with unified text and image masking

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:31.264667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.037690Z digest=sha256:62e56ac27d8c785b5a8c5faa121cb0ec4554a5f843e799fefa8bb5fd88482037

Observation 464c226c-995c-479a-94e3-24e3fd79c8cc · outbound

This paper cites TrustLLM: Trustworthiness in Large Language Models.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness TrustLLM: Trustworthiness in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.102929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.102929Z digest=sha256:908102f84c092b9badcfee98fd904780c35eb351f0a92e90e63aa1aed35e43a6

Observation 4fc0f8b1-a59b-4dff-91f2-a8596553d501 · outbound

This paper cites Icdar2019 competition on scanned receipt ocr and information extraction.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Icdar2019 competition on scanned receipt ocr and information extraction

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:31.156302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.195342Z digest=sha256:2f4a9465541d9e2b1887486b07062143b2811602b7df7a9cc5ed455ed7bddbc9

Observation 33ce8a36-0753-4860-9bcc-56f1e33e89bd · outbound

This paper cites From image to language: A critical analysis of visual question answering (vqa) approaches, challenges, and opportunities.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness From image to language: A critical analysis of visual question answering (vqa) approaches, challenges, and opportunities

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:31.094217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.220974Z digest=sha256:a409315edcfebaeb310602615590f79cc98e219409362c6d6bbff2f0c1a3946e

Observation c289804d-c757-4bc5-8362-939f304538a0 · outbound

This paper cites Funsd: A dataset for form understanding in noisy scanned documents.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Funsd: A dataset for form understanding in noisy scanned documents

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:31.077837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.240067Z digest=sha256:e748fc29db895147848b0225b1a520f476c627441c908f4e43fdff790a3b172e

Observation c902eb2d-3712-4613-b96b-5fc156ac7031 · outbound

This paper cites Ocr-free document understanding transformer.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Ocr-free document understanding transformer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:31.061211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.260382Z digest=sha256:7a0c233d59fdfac9f9765afac9bc6398469fb475265041666575feb85be6f455

Observation 095cf3f5-8ab4-44e6-9364-1bffa057965d · outbound

This paper cites Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language Models.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.276604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.276604Z digest=sha256:7932656ade9d73d8c9550cab6f8d1486b4507f6491318f4cf1dadd5c8e58c6ca

Observation 91cd93d8-0d02-44e7-96ff-8281be9747c7 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness LLaVA-OneVision: Easy Visual Task Transfer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.281806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.281806Z digest=sha256:89f34e3a55ef7d3c1c3cf1f47f99b240314eeed471cab704616a88298f03a888

Observation 6bcf163b-ccfa-429d-840b-cf3c19746ea1 · outbound

This paper cites Show, attend and read: A simple and strong baseline for irregular text recognition.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Show, attend and read: A simple and strong baseline for irregular text recognition

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.906380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.287419Z digest=sha256:b0cabe59ceb9c303127406e610cb38013a456b880b72f2e27654ba28d905173e

Observation 14374c31-e293-4bdf-9370-5253ac943455 · outbound

This paper cites Trocr: Transformer-based optical character recognition with pre-trained models.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Trocr: Transformer-based optical character recognition with pre-trained models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.869327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.292071Z digest=sha256:a1ba101e4311a3ec92c0bab8216a77e4d85d35b7d01d74adcc433b4ef6ac0ec8

Observation 0a23538a-cf0a-4bb4-8c83-487443b8c7a2 · outbound

This paper cites Real-time scene text detection with differentiable binarization.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Real-time scene text detection with differentiable binarization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.854445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.297851Z digest=sha256:7e06bff461e68f86ff88bb8cc2dcd204aed85dfa0456622d227314edda143d75

Observation cea610ae-b9d6-4768-9db0-8826fbc32897 · outbound

This paper cites DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.303009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.303009Z digest=sha256:9eaa87c1b219caa5b698e6cf915f3056143452dc973476d56dbfce5d0202d2a6

Observation 04012c68-cb70-46e2-bc58-f28a534908ac · outbound

This paper cites DoRA: Weight-Decomposed Low-Rank Adaptation.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness DoRA: Weight-Decomposed Low-Rank Adaptation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.308712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.308712Z digest=sha256:e767fbf7ef3f41291b860dfc226ad89ded298d17a7ec48b9b82d804d8ea18345

Observation 24918f95-f5ac-4e13-a9b6-509d24b8efd8 · outbound

This paper cites A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.314172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.314172Z digest=sha256:62c0e3677442cad20703dfc03ff58c7cc9dca7a20b0482ffa373a2445fba1026

Observation 41adb064-175b-4bf1-9bd7-8b884ddd6f79 · outbound

This paper cites Master: Multi-aspect non-local network for scene text recognition.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Master: Multi-aspect non-local network for scene text recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.832727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.319771Z digest=sha256:4904eb5c437857235563b3e8b6d42cd2fa8717cd6851f2ae209459d220e789cf

Observation 3c7031a7-8f4a-41f9-9247-c52d6a116b24 · outbound

This paper cites Layoutllm: Layout instruction tuning with large language models for document understanding.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Layoutllm: Layout instruction tuning with large language models for document understanding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.666353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.324860Z digest=sha256:8d4441ee3ca2c81c02606d661f6ce7cb563e58a0b90c645d732fac072c691872

Observation c55309fc-07ea-4dcf-ae19-6aef6a5878c7 · outbound

This paper cites MaskOCR: Text Recognition with Masked Encoder-Decoder Pretraining.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness MaskOCR: Text Recognition with Masked Encoder-Decoder Pretraining

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.397549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.397549Z digest=sha256:fe3fd9c550384a46257592efe63a3cdd155004e3a633b33366891d5b847276fb

Observation e8a5b922-66ea-488d-a05f-58ba5331d102 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Docvqa: A dataset for vqa on document images

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.484467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.484467Z digest=sha256:b62fe9325bbe7d733b021eb40ce5d6d64a0c7ca83e5ea0d58d448b74de1a640b

Observation db6c78e2-d984-4198-a42c-2a2ef5d21ec1 · outbound

This paper cites Cord: a consolidated receipt dataset for post-ocr parsing.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Cord: a consolidated receipt dataset for post-ocr parsing

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.620032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.620032Z digest=sha256:b8e540b639b435827cc9e508d2d54bc10cf3a62ed910008607d9b09d43425d25

Observation 046668fb-9f94-48ed-900a-fea9b353c07c · outbound

This paper cites Generalized intersection over union: A metric and a loss for bounding box regression.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Generalized intersection over union: A metric and a loss for bounding box regression

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.486734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.711859Z digest=sha256:55ad4c928c16e829ab1ae147d84562e2e8d2335fae9e51ba6e1abb03cb398feb

Observation 0dff8af7-c726-4a5c-b031-9daa5ef25106 · outbound

This paper cites An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.461824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.764976Z digest=sha256:12e39acf0e6377cdd7bedea81bc025e3924bf6a82a2bb3675a4ffab633aa8c91

Observation 09931be9-998f-4adf-8e53-60553123597a · outbound

This paper cites Instructdoc: A dataset for zero-shot generalization of visual document understanding with instructions.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Instructdoc: A dataset for zero-shot generalization of visual document understanding with instructions

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.437769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.792352Z digest=sha256:5ccb6793fcf22cc93437bb92e2d5e810a0a4061a7ee01cb2bd20390bf9efe5c4

Observation 67b0da32-b059-4f5f-a8f7-adba54f3ba71 · outbound

This paper cites Unifying vision, text, and layout for universal document processing.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Unifying vision, text, and layout for universal document processing

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.390476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.825256Z digest=sha256:2ce2c1b4dbc344b033b01766c429b00090138a5cf81fc5a0b7f57bdc265dc2b1

Observation 0138eb19-e9a8-4353-8477-4af778f08be8 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.848452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.848452Z digest=sha256:eb92df7e917a8543345d7bbf02ee39b3878448d59b1362f77f5e1af40151c950

Observation ec38c8da-6808-4311-8cf9-b1aed03653b8 · outbound

This paper cites Omniparser: A unified framework for text spotting key information extraction and table recognition.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Omniparser: A unified framework for text spotting key information extraction and table recognition

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T10:13:30.348538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-12T10:13:29.858078Z digest=sha256:68b96bae0e61314ed85d3313417acf149d8daca4658c7c84c006f25f8bade38f

Observation 56d10951-a55a-4939-ab07-6aa3e32019f4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.869788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.869788Z digest=sha256:36400a2143fd6250f2b33e0e9a04a3350afc8be5903eaa1300ca15dda6ad5352

Observation ba2aeb0b-fcf8-4247-bb6e-ffd64ea540cf · outbound

This paper cites Layout and Task Aware Instruction Prompt for Zero-shot Document Image Question Answering.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Layout and Task Aware Instruction Prompt for Zero-shot Document Image Question Answering

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.874227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.874227Z digest=sha256:026fb1ddec61c96925888a1cfc314a2a212d2f4513f1489fd12e81b6bd4e7462

Observation 938f48a7-1a9f-4219-b363-e08acc3aa3d5 · outbound

This paper cites A normalized levenshtein distance metric.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness A normalized levenshtein distance metric

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.878827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.878827Z digest=sha256:79048db9bd5093139b57960e6f409875d279f2bd87b72c5de55c64ad75e1d103

Observation d140bdda-db4f-4397-b489-6a1a58f4cafe · outbound

This paper cites MixNet: Toward Accurate Detection of Challenging Scene Text in the Wild.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness MixNet: Toward Accurate Detection of Challenging Scene Text in the Wild

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.882984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.882984Z digest=sha256:e7ea58e31edb72abda9cbc6c4919928ceb9f86675a8c7c9dc50b4b99b617a487

Observation 7edc6aaf-2860-41f6-8347-58a8f0931ca9 · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T10:13:29.887857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T10:13:29.887857Z digest=sha256:0573e23e5dfbf29be930351758627e7360b1c1015c6c6ca3f6f052bd08b04c07

Pith citing papers

Observation b8fcff72-8148-4bbc-8735-301968efb62e · inbound

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering cites this paper.

Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:53.544005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:07:53.544005Z digest=sha256:470ecec3641b6209e0eecf48bffb1b1a4df315d4cfc9d705bf56a04e963e13b0

Observation 8e5826c9-dd94-46cc-8d8c-4fe1a8678444 · inbound

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA cites this paper.

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:44:02.459977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T04:41:55.030673Z digest=sha256:65aad5b0b0368b6b48c457129467a6c54b0121785cd3d33cb0327e9e5a2fd8dd

Observation 13c8431a-09e6-43fe-bf40-f1ecb82f34a6 · inbound

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment cites this paper.

Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:28.323502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:25:28.323502Z digest=sha256:e4250fada954549755848e42cf72802185f452802f6d2352b1ef3a1863f5f362