Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:42:14.915935Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2412.08746.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:42:14.915935Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-20T21:55:02.656863Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T21:59:06.400938Z
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0d95e635-eb90-46bc-8478-43366a9a697e · outbound
DocVLM: Make Your VLM an Efficient Reader Sequence-to-sequence contrastive learning for text recogni- tion
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ce335f3e-c1c6-4b33-a514-a18942a47dd9 · outbound
DocVLM: Make Your VLM an Efficient Reader Multimodal Semi-Supervised Learning for Text Recognition
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a83da128-c292-4c5a-82f0-a0123d3c23a5 · outbound
DocVLM: Make Your VLM an Efficient Reader Clipter: Looking at the bigger picture in scene text recogni- tion
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 97adfa2a-0880-4d59-84bf-f52aa06e0634 · outbound
DocVLM: Make Your VLM an Efficient Reader Visfocus: Prompt- guided vision encoders for ocr-free dense document under- standing
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0dbb5cf3-6f5c-4557-bb3e-4aa26e03ea85 · outbound
DocVLM: Make Your VLM an Efficient Reader Flamingo: a visual language model for few-shot learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abc368ed-db23-432e-b4c8-6a799a689b42 · outbound
DocVLM: Make Your VLM an Efficient Reader Docformer: End-to-end transformer for document understanding
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 33997b18-557c-491d-bfe4-53f9caf214f8 · outbound
DocVLM: Make Your VLM an Efficient Reader Docformerv2: Local features for document understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 603ac7e8-1290-4fd2-acee-b26cfdd395dd · outbound
DocVLM: Make Your VLM an Efficient Reader ScreenAI: A Vision-Language Model for UI and Infographics Understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86291faa-31e1-4997-9aed-7da95e13ebe8 · outbound
DocVLM: Make Your VLM an Efficient Reader Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4053255-729c-4b5c-9baa-ecd98b1ba15f · outbound
DocVLM: Make Your VLM an Efficient Reader PaliGemma: A versatile 3B VLM for transfer
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e5bfb32-75f0-4d42-90f3-71b82931c738 · outbound
DocVLM: Make Your VLM an Efficient Reader Scene text visual question answering
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 680bee12-b608-4c91-9c1a-8a1c292b6423 · outbound
DocVLM: Make Your VLM an Efficient Reader Latr: Layout-aware transformer for scene-text vqa
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a19bad3c-e51f-4dd0-8e1b-9e85c7880300 · outbound
DocVLM: Make Your VLM an Efficient Reader OCR-IDL: OCR Annotations for Industry Document Library Dataset
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b3582b1-5aba-429e-9051-a6e8b13c226c · outbound
DocVLM: Make Your VLM an Efficient Reader Gram: Global reasoning for multi-page vqa
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 263231ec-21af-428e-bcf0-b8591b249bcc · outbound
DocVLM: Make Your VLM an Efficient Reader Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70b2f9db-b7f6-457e-9d1f-2638aeb6eebb · outbound
DocVLM: Make Your VLM an Efficient Reader PaLI-X: On Scaling up a Multilingual Vision and Language Model
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0db02ade-4d5c-4fc9-80c0-dd8a8cc24eb9 · outbound
DocVLM: Make Your VLM an Efficient Reader PaLI-3 Vision Language Models: Smaller, Faster, Stronger
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3809b3e9-1c73-4ba8-b5c8-1daba241f17c · outbound
DocVLM: Make Your VLM an Efficient Reader Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e34e2c5d-740a-430a-bb97-3b315914d230 · outbound
DocVLM: Make Your VLM an Efficient Reader InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8a84652-7a5b-4304-b070-6d2d8702753e · outbound
DocVLM: Make Your VLM an Efficient Reader InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d08ab50-a9f8-47a5-a6ea-1497d6abc638 · outbound
DocVLM: Make Your VLM an Efficient Reader Dtrocr: Decoder-only transformer for op- tical character recognition
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8e22a9ff-be1b-49c4-b305-6c5d0c8262c3 · outbound
DocVLM: Make Your VLM an Efficient Reader Towards models that can see and read
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bb305f88-a348-4121-8b18-6bdf6dbe89d1 · outbound
DocVLM: Make Your VLM an Efficient Reader Question aware vision transformer for multimodal reasoning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 420b070a-ff4b-4341-bc40-f9ab7456889e · outbound
DocVLM: Make Your VLM an Efficient Reader Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 74329bad-5407-48c5-8423-7176f78b40f2 · outbound
DocVLM: Make Your VLM an Efficient Reader Funsd: A dataset for form understanding in noisy scanned documents
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c256a160-35cd-4c4f-85dc-b43b8871cd0f · outbound
DocVLM: Make Your VLM an Efficient Reader M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 06682577-5388-4b29-b7fa-a806ecdc8dc3 · outbound
DocVLM: Make Your VLM an Efficient Reader mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33ef0390-6e67-4dfa-9864-6e3c4f3dac4d · outbound
DocVLM: Make Your VLM an Efficient Reader Towards unified scene text spotting based on sequence generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e2e9d1b4-b0b2-40de-981a-93cbddac54b1 · outbound
DocVLM: Make Your VLM an Efficient Reader OCR-free Document Understanding Transformer
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7185aa3e-4c12-40d5-a219-ce5e089d5ed4 · outbound
DocVLM: Make Your VLM an Efficient Reader What matters when building vision-language models?
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cccefa27-388d-4548-a17e-43f726d64b16 · outbound
DocVLM: Make Your VLM an Efficient Reader LLaVA-OneVision: Easy Visual Task Transfer
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 808b0bc4-6e05-4254-bdce-04d498a0186a · outbound
DocVLM: Make Your VLM an Efficient Reader Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0873528-fedb-49f4-88f7-8c9689af5dca · outbound
DocVLM: Make Your VLM an Efficient Reader TokenPacker: Efficient Visual Projector for Multimodal LLM
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1cb8b79-8268-4011-8271-c5f017051d93 · outbound
DocVLM: Make Your VLM an Efficient Reader Scatter: selective con- text attentional scene text recognizer
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9fdad15d-55ce-4431-9fe5-4d0bf66f134d · outbound
DocVLM: Make Your VLM an Efficient Reader Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9272f10-53a6-48da-b75c-b78946a74760 · outbound
DocVLM: Make Your VLM an Efficient Reader Visual instruction tuning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd5f3872-6ef9-4f77-8e34-ccbb466a6e86 · outbound
DocVLM: Make Your VLM an Efficient Reader Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4be20eb9-54f9-4a2e-8cb7-53b5e0f58641 · outbound
DocVLM: Make Your VLM an Efficient Reader ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83b1b8fb-4943-4c57-931b-5ba6817448e9 · outbound
DocVLM: Make Your VLM an Efficient Reader Docvqa: A dataset for vqa on document images
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 49244a03-65ff-401b-8ff1-1f842a332346 · outbound
DocVLM: Make Your VLM an Efficient Reader Infographicvqa
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 19674d70-1e3f-49a8-88e9-4b1ee35b7cd7 · outbound
DocVLM: Make Your VLM an Efficient Reader Ocr-vqa: Visual question answering by reading text in images
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aec45507-a783-4669-86b1-4295e6e6e4fd · outbound
DocVLM: Make Your VLM an Efficient Reader Textadain: Paying attention to shortcut learning in text recognizers
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation eca49a30-8850-46bd-a733-9f620a43f6f5 · outbound
DocVLM: Make Your VLM an Efficient Reader Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1acf826-880f-4844-87cb-0403900ba6dd · outbound
DocVLM: Make Your VLM an Efficient Reader Exploring the limits of transfer learning with a unified text-to-text transformer
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4c32e6e-f4a5-4dfa-b566-d7031b1c6888 · outbound
DocVLM: Make Your VLM an Efficient Reader GLASS: Global to Local Attention for Scene-Text Spotting
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22912a0c-9d61-4879-b8fe-13698debcdd3 · outbound
DocVLM: Make Your VLM an Efficient Reader Textcaps: a dataset for image caption- ing with reading comprehension
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d30a0a9-8f37-44cb-99c9-20d469c48152 · outbound
DocVLM: Make Your VLM an Efficient Reader Towards vqa models that can read
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e2a6e898-08b5-4869-b4ea-79fc57b94d29 · outbound
DocVLM: Make Your VLM an Efficient Reader Instructdoc: A dataset for zero-shot gener- alization of visual document understanding with instructions
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4fc6056f-ba81-4d87-9163-bd3fbb4aa3c4 · outbound
DocVLM: Make Your VLM an Efficient Reader Hi- erarchical multimodal transformers for multipage docvqa
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f2f38ca2-8ee0-429e-a95b-d9d29bfbb79a · outbound
DocVLM: Make Your VLM an Efficient Reader Document understanding dataset and evaluation (dude)
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8a1ab35e-680a-4395-8b0d-539f5af5fb83 · outbound
DocVLM: Make Your VLM an Efficient Reader DocLLM: A layout-aware generative language model for multimodal document understanding
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6e054a9-6d96-4299-8b1f-da5b50f2d127 · outbound
DocVLM: Make Your VLM an Efficient Reader Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c96d11ff-42d2-4c2a-9ca8-414d0cfd81f1 · outbound
DocVLM: Make Your VLM an Efficient Reader Layout and Task Aware Instruction Prompt for Zero-shot Document Image Question Answering
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b02716ba-9456-4415-9e20-88f23408f67f · outbound
DocVLM: Make Your VLM an Efficient Reader Layoutlm: Pre-training of text and layout for document image understanding
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9f687888-e6cb-4ea7-8640-113138fd9b41 · outbound
DocVLM: Make Your VLM an Efficient Reader UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9cca9cb-c68d-4cc0-8bb8-54592e6afc9e · outbound
DocVLM: Make Your VLM an Efficient Reader Dptext-detr: Towards better scene text detection with dynamic points in transformer
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a6f9ff17-c632-449a-b020-fb0fb9f84129 · outbound
DocVLM: Make Your VLM an Efficient Reader Deepsolo: Let transformer decoder with explicit points solo for text spot- ting
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4d194bfd-b861-4324-b4ef-628cc0342fc5 · outbound
DocVLM: Make Your VLM an Efficient Reader mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 381c7c14-7bc2-468b-adad-b45847ef488a · outbound
DocVLM: Make Your VLM an Efficient Reader InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3afa750-7b12-474b-ae8e-56f25805e8af · outbound
DocVLM: Make Your VLM an Efficient Reader MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23c2e6cd-2f15-4de5-b2ee-35443e785d9d · outbound
DocVLM: Make Your VLM an Efficient Reader Towards complex doc- ument understanding by discrete reasoning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c79542d8-7131-404a-813b-9090a688b5c9 · outbound
DocVLM: Make Your VLM an Efficient Reader The encoder is initial- ized with pretrained weights from DocFormerV2, which was pretrained on the Industry Document Library (IDL) dataset [13]
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 83f55883-6cbf-4800-a317-78e8c3ccad98 · inbound
Operationalizing Document AI: A Microservice Architecture for OCR and LLM Pipelines in Production DocVLM: Make Your VLM an Efficient Reader
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.