Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:59:07.601054Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2507.09531.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T17:59:07.601054Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 29dfd4c0-64c6-4a2c-acdd-19fef097dacb · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Docformer: End-to-end transformer for document understanding
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5508c845-f007-4386-b509-bdb252ee4268 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Qwen-vl: A versatile vision-language model for un- derstanding, localization
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 39ece4dc-3400-49d9-a18a-1fba010b1dc6 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Due: End-to-end document understand- ing benchmark
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e0ffee70-b30d-4c9a-a799-f5fae76f8f5d · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be0e5316-1b31-4aa0-94eb-a6704667ce56 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Gonzalez, Ion Stoica, and Eric P
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ad2367e2-b82d-45f9-8574-151d13ee6d98 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Instructblip: towards general- purpose vision-language models with instruction tuning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2562b1e0-86bb-4dd2-b276-104e062ac7ab · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Docpedia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 72ea372f-a773-49ab-b868-f35ea1b0ddd1 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Unidoc: Unified pretraining framework for document understanding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e1a838da-28c7-4965-8f29-5a4561e6547e · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Deep residual learning for image recognition
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b584473d-7079-4be1-ace2-70cadbd24256 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Gaussian Error Linear Units (GELUs)
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fec0b0f-ce6d-4e54-ad0f-301f5e9fcad5 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Cogagent: A visual language model for gui agents
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 33ee729d-cb83-4a84-8245-2eb18fe80198 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization SciCap: Generating captions for scientific figures
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d7737523-5de0-43c3-a758-44fe24b45c1a · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization mPLUG-DocOwl 1.5: Unified structure learning for OCR- free document understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9bff7f50-ba8c-4efd-80dd-554c471ec62a · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Lora: Low-rank adaptation of large language models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6606b3e1-991a-42ff-b0e2-62480fe7f54c · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Layoutlmv3: Pre-training for document ai with unified text and image masking
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 85147d98-064b-4328-8385-1b3d0f26df68 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Icdar2019 compe- tition on scanned receipt ocr and information extraction
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bc7d70da-177f-46f5-99f3-9c8024905df1 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Funsd: A dataset for form understanding in noisy scanned documents
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 320e75f1-9534-4125-b13f-1c8c709c3879 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization A diagram is worth a dozen images
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfc00b05-2fac-41d0-b892-0692a689a5ee · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization OCR-free Document Understanding Transformer
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6689e449-4978-47f0-ae20-f53b79396df4 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e7f6cae0-596e-4cba-adb2-cdd848a8e768 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization DocBank: A bench- mark dataset for document layout analysis
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9729aeed-e6c8-4829-95ab-5a0d911f373b · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Selfdoc: Self-supervised document representation learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c7a26764-3cc7-479d-8571-a0fc72c79c92 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Mon- key: Image resolution and text label are important things for large multi-modal models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e2683014-4149-4068-9dce-8ebb7e5ef479 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ad7ffd0-f94a-426d-80db-5125b053fac6 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Microsoft coco: Common objects in context
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d80c291-eed3-44fe-9945-e8fa849f5100 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Feature pyra- mid networks for object detection
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7fb894e-2d7d-48e7-8cc2-7cd6a526d8ac · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Visual instruction tuning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02842439-b046-4944-b9c5-5a46b0a4d168 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Improved baselines with visual instruction tuning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cc6f6b2f-255c-40f5-b1df-25125e92583b · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 691dc397-de21-4031-8702-cf906a213fb4 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Swin transformer v2: Scaling up capacity and resolution
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f052739d-9cc9-4a47-94f9-cabd4510cc10 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization SGDR: Stochastic Gradient Descent with Warm Restarts
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7544abc3-553f-4d8b-90f1-89d96377dff3 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Decoupled Weight Decay Regularization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90a8901b-f628-45d3-a1e1-3cf2a4e7b441 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40e38e9f-cae0-47d4-9c06-2aae302fb3f7 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization DocVQA: A Dataset for VQA on Document Images
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33e1d906-ab40-46a0-bc34-7c2a5cd57de4 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Azure Cognitive Services: Optical Char- acter Recognition (OCR)
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bc0fdc9f-f52d-4fb0-a1f5-a8e1a0ee34e9 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Rectified linear units im- prove restricted boltzmann machines
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dc51e973-6da0-455a-a2b2-dc51d816a9c2 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Cord: a con- solidated receipt dataset for post-ocr parsing
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0534d87b-bb01-452c-bb25-29b97a652f07 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Automatic differentiation in pytorch
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 400d844d-0b27-4646-8b4a-c68848708a74 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Doclaynet: A large human- annotated dataset for document-layout segmentation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 519d9a20-7fff-4027-8483-af81641f1a3a · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Going full-tilt boogie on document understanding with text-image-layout transformer
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5761f58d-16d5-448b-9b94-5b221ebb6dd6 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5ef67d07-a8dc-41dc-9d26-cb56b53ddcc9 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Faster r-cnn: Towards real-time object detection with region proposal networks
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9c86612-5a36-4025-a683-38c90fd824d6 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Deep Learning based Key Information Extraction from Business Documents: Systematic Literature Review
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0450fa45-eb9f-42af-997c-a167f0ea5ae5 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Docile benchmark for document information localization and extraction
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4f4ff3ee-551a-480a-b2f8-3c623614cef3 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Towards vqa models that can read
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7927cf82-2d59-4f37-a184-cef2e104c169 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Spatial Dual-Modality Graph Reasoning for Key Information Extraction
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abe05721-78e1-4dea-a77f-29a5e9be424e · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Instructdoc: A dataset for zero-shot general- ization of visual document understanding with instructions
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1fda786d-7401-4228-913d-144790c9ec8e · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Docllm: A layout-aware gener- ative language model for multimodal document understand- ing
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 93db90e9-a191-4f32-86b2-5df0d338957e · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Vision-enhanced semantic entity recognition in document images via visually-asymmetric consistency learning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9212ef6d-d603-452e-9f5f-855e5e9601d3 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6292d008-1c23-4547-8c70-31718d9802ac · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Layoutlm: Pre-training of text and layout for document image understanding
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 86758d95-0187-4e1b-bf58-47ca6d7d7f53 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Layoutlmv2: Multi-modal pre-training for visually-rich document understanding
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2454a794-357c-430c-9d0d-367ccdaf14d4 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c7ea86d-7888-4a3b-a0d7-2bf2eeb9d5ad · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization UReader: Universal OCR-free visually-situated language understand- ing with multimodal large language model
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 249d394c-c77b-4bbf-8550-f94411c7e77b · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization By my eyes: Grounding multimodal large language models with sensor data via vi- sual prompting
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cd4dc51e-b3e7-4372-80dd-ac741b1daf00 · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1a888602-e291-42c8-9625-1461a92fcc8a · outbound
VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.