Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:13:29.887857Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 3 inbound Pith citation observations for arXiv:2412.00151.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T10:13:29.887857Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:07:53.544005Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-17T04:44:02.457882Z
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e0d80dc3-aaf1-4e20-a0dd-f274cd6183cb · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bcb9371-e82c-4bc1-8af1-7e0b7714cf15 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Pixtral 12B
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f71ecc7-3cfa-419d-b13c-962ce18621af · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Vision transformer for fast and efficient scene text recognition
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 233e36ff-921b-4862-861c-7d5b6cd666e7 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Scene text recognition with permuted autoregressive sequence models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fdf3778b-c7ff-4cc8-9bd0-9c1251f9c726 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness FAST: Faster Arbitrarily-Shaped Text Detector with Minimalist Kernel Representation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e4be6ac-9f30-4f9f-be78-d302dcc71de1 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4444bbda-73bc-4e82-872f-3be4203913e7 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83f3f23e-022b-4587-8287-c376b7ab17cd · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 709d6934-67e5-454f-b908-494e30aa8bd3 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Qlora: Efficient finetuning of quantized llms
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ece1e68a-f183-4f67-af7a-b307804581c2 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6ff43f4-b7f9-4c69-8ed1-f37a2a170e15 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Dtrocr: Decoder-only transformer for optical character recognition
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f1f9abc0-e275-4b8e-a3e3-aae723b5c98d · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Retrieval-Augmented Generation for Large Language Models: A Survey
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7da7fed-e7a0-485d-b614-a55109a7b570 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness LoRA+: Efficient Low Rank Adaptation of Large Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cea5e5f-0040-4241-a0cd-613349378ddc · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Icl-d3ie: In-context learning with diverse demonstrations updating for document information extraction
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 46c45ae0-8534-46bd-8d5c-3983522dec5a · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness LoRA: Low-Rank Adaptation of Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd3eeaa5-026f-4530-98cc-0a5118ec88b8 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Layoutlmv3: Pre-training for document ai with unified text and image masking
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 464c226c-995c-479a-94e3-24e3fd79c8cc · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness TrustLLM: Trustworthiness in Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fc0f8b1-a59b-4dff-91f2-a8596553d501 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Icdar2019 competition on scanned receipt ocr and information extraction
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 33ce8a36-0753-4860-9bcc-56f1e33e89bd · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness From image to language: A critical analysis of visual question answering (vqa) approaches, challenges, and opportunities
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c289804d-c757-4bc5-8362-939f304538a0 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Funsd: A dataset for form understanding in noisy scanned documents
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c902eb2d-3712-4613-b96b-5fc156ac7031 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Ocr-free document understanding transformer
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 095cf3f5-8ab4-44e6-9364-1bffa057965d · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Visually-Situated Natural Language Understanding with Contrastive Reading Model and Frozen Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91cd93d8-0d02-44e7-96ff-8281be9747c7 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness LLaVA-OneVision: Easy Visual Task Transfer
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bcf163b-ccfa-429d-840b-cf3c19746ea1 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Show, attend and read: A simple and strong baseline for irregular text recognition
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 14374c31-e293-4bdf-9370-5253ac943455 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Trocr: Transformer-based optical character recognition with pre-trained models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0a23538a-cf0a-4bb4-8c83-487443b8c7a2 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Real-time scene text detection with differentiable binarization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation cea610ae-b9d6-4768-9db0-8826fbc32897 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04012c68-cb70-46e2-bc58-f28a534908ac · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness DoRA: Weight-Decomposed Low-Rank Adaptation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24918f95-f5ac-4e13-a9b6-509d24b8efd8 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41adb064-175b-4bf1-9bd7-8b884ddd6f79 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Master: Multi-aspect non-local network for scene text recognition
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3c7031a7-8f4a-41f9-9247-c52d6a116b24 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Layoutllm: Layout instruction tuning with large language models for document understanding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c55309fc-07ea-4dcf-ae19-6aef6a5878c7 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness MaskOCR: Text Recognition with Masked Encoder-Decoder Pretraining
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8a5b922-66ea-488d-a05f-58ba5331d102 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Docvqa: A dataset for vqa on document images
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db6c78e2-d984-4198-a42c-2a2ef5d21ec1 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Cord: a consolidated receipt dataset for post-ocr parsing
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 046668fb-9f94-48ed-900a-fea9b353c07c · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Generalized intersection over union: A metric and a loss for bounding box regression
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0dff8af7-c726-4a5c-b031-9daa5ef25106 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 09931be9-998f-4adf-8e53-60553123597a · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Instructdoc: A dataset for zero-shot generalization of visual document understanding with instructions
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 67b0da32-b059-4f5f-a8f7-adba54f3ba71 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Unifying vision, text, and layout for universal document processing
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0138eb19-e9a8-4353-8477-4af778f08be8 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec38c8da-6808-4311-8cf9-b1aed03653b8 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Omniparser: A unified framework for text spotting key information extraction and table recognition
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 56d10951-a55a-4939-ab07-6aa3e32019f4 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba2aeb0b-fcf8-4247-bb6e-ffd64ea540cf · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness Layout and Task Aware Instruction Prompt for Zero-shot Document Image Question Answering
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 938f48a7-1a9f-4219-b363-e08acc3aa3d5 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness A normalized levenshtein distance metric
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d140bdda-db4f-4397-b489-6a1a58f4cafe · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness MixNet: Toward Accurate Detection of Challenging Scene Text in the Wild
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7edc6aaf-2860-41f6-8347-58a8f0931ca9 · outbound
DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8fcff72-8148-4bbc-8735-301968efb62e · inbound
Spatially Grounded Explanations in Vision Language Models for Document Visual Question Answering DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e5826c9-dd94-46cc-8d8c-4fe1a8678444 · inbound
DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 13c8431a-09e6-43fe-bf40-f1ecb82f34a6 · inbound
Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering via Reasoning-Free Alignment DLaVA: Document Language and Vision Assistant for Answer Localization with Enhanced Interpretability and Trustworthiness
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.