Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:34:31.918078Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 6 inbound Pith citation observations for arXiv:2411.17125.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:34:31.918078Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:09:11.055077Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T23:35:07.388169Z
64 of 64 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 26bfbaeb-8f0a-4d97-81af-e54ea2c53471 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24a4114f-9514-43c0-99d1-d27c2a3f1906 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Jawahar, Ernest Valveny, and Dimos- thenis Karatzas
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3bdda016-7e96-44d9-82cc-85788fd7cd7a · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Textocr-gpt4v
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 17f76eeb-cb04-45da-ac8d-ea8681731f84 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bd12d6c-c3d4-4aa0-aeb2-f7bcec147bfb · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Tabfact : A large-scale dataset for table-based fact verification
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6a3282b8-7a6b-4543-9a1f-523f3ca78d66 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11801deb-4f35-47ee-946e-a8a415b08cbe · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ae2b981-6d89-453b-96ca-01ddd952001c · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5be5ab8-9bc1-44a1-add0-7135206b6aa0 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Hitab: A hierarchical table dataset for question an- swering and natural language generation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0833509a-d22a-43bd-b5b2-dc5b3badfb5a · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86c18ed8-c0ac-4dd9-8bae-979d92f363c7 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Pymupdf: Python bindings for mupdf
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 02369ccb-e3bd-494a-a328-777d4ed14c00 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff194cb4-9bee-4e45-ab8c-d9081b52d5a2 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c2d8526-7b32-4d4b-ab24-bd8dbacfbb5b · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring mplug- docowl2: High-resolution compressing for ocr-free multi- page document understanding, 2024
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c4b5a9a1-23d7-4426-9464-4172398ab6e7 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring GPT-4o System Card
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 685b00ea-5278-4467-b538-07ddaa96579e · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Dvqa: Understanding data visualizations via ques- tion answering
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bf19eda8-5938-4e02-9bbe-f4aa97156073 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Fig- ureqa: An annotated figure dataset for visual reasoning,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abc404f9-3788-4835-88f4-690e59522ef0 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring A diagram is worth a dozen images
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3201579a-ddad-4b98-bd1a-cf897857c69a · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Are you smarter than a sixth grader? textbook question answer- ing for multimodal machine comprehension
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8254ade0-206f-4b5e-bed7-4b9587714bb1 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Ocr-free document understanding transformer
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bd41f984-f577-42d4-a053-51f817060538 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 233f6a0a-37ed-4caf-b440-9f7b4dbe29bb · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f8c0dc91-a888-4c5e-8823-ec4ddd7802b6 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Focus Anywhere for Fine-grained Multi-page Document Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82070d9c-33c0-4d3d-a63d-994a987b1975 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b86853e7-b33e-4e6d-a01f-1ae64ff2d84f · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a49c185-310c-466b-9d49-ddf16efa2cb9 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Visual instruction tuning, 2023
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf4bc3c0-0fc8-4a09-808e-3e2358f57ebe · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Textmonkey: An ocr-free large multimodal model for understanding document, 2024
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 55824c0c-9478-463f-8219-ebe694609b52 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring KOSMOS-2.5: A Multimodal Literate Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b46eb2d2-7fa6-4847-8a65-a189fd7ff92f · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring The iam-database: an english sentence database for offline handwriting recognition
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b971a719-8ed5-4cd5-801f-d4ba6aa2f63c · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring ChartQA: A benchmark for question answering about charts with visual and logical reasoning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cfb35cac-fa5c-485d-8c02-57a3c615389e · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Chartqa: A benchmark for question an- swering about charts with visual and logical reasoning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4d854a3b-a335-4c5b-b400-fb0d838884a2 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fc59c0a9-032d-41cc-917b-9dcc68d1af8a · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ff135251-6a65-4408-b668-52e9ca751260 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Mishra, K
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f8b951c5-e9b6-4b9e-ad33-16912428e8d7 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Chart-to-text: Generat- ing natural language descriptions for charts by adapting the transformer model, 2020
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 73c958f6-3054-4c68-aedd-73337ca00cd2 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Compositional semantic parsing on semi-structured tables
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bcb6e8df-9dc2-4b8c-85fc-5b4f0032ac27 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Kosmos-2: Ground- ing multimodal large language models to the world
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4006fc3-7d27-4cd1-94da-e727d5633d65 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 139bd0b3-a96d-496f-9f22-13e542788837 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Textcaps: a dataset for image caption- ingwith reading comprehension
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6aee4a9f-96da-4d56-91db-0d9ac003b27e · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Towards vqa models that can read
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dd376bc9-496b-41e7-9f7c-a4a872091e98 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Kleister: Key in- formation extraction datasets involving long documents with complex layouts
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 53f13a80-5f57-42a4-b9d7-25d681224026 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Deepform: Understand structured docu- ments at scale
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b6e509b7-96b3-4561-9974-20b284045ab9 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Vi- sualmrc: Machine reading comprehension on document im- ages
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 16a1f5ae-9a57-492d-ab18-c1ea2f67bf5d · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Tang, Angie Boggust, and Arvind Satyanarayan
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a88d28f-3ed9-43c3-9fab-4cf55b9231c2 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c97f8457-39f0-4a02-b8e9-5a0f22fd0dcb · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Cc-main-2021-31-pdf- untruncated
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 73d5a102-9789-401d-ad47-2de3a29363ea · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 002e889c-ada0-4abe-8e71-2fd052a4fbfe · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Lawrence Zitnick, and Devi Parikh
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 833c0e39-612c-45fd-992d-e3ab4ab6ccd2 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Screen2words: Automatic mobile ui summarization with multimodal learning, 2021
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a6ae5d35-1d18-43a3-8b4e-ec6e9f6e336c · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Mineru: An open-source solution for precise document content extrac- tion, 2024
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fb01f736-d8c8-492b-8506-5086074fa2e6 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 03780143-2616-445c-ba74-cfe102edf024 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91f13e26-c975-4900-a338-4ae045535b86 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring wendlerc/renderedtext, 2023
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 39990f8d-4f61-4327-acb1-f0f1310396da · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 397fcff5-cc26-4214-9c8a-d9b81bec51da · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Canvasvae: Learning to generate vector graphic documents
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f46af81-460d-4109-9a0e-016a452736b2 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Qwen2 Technical Report
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e472e1d7-8b69-4fb1-8f6c-ecbb9889acab · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1852f1e3-5581-4968-921f-5172e00bd02c · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Ferret: Refer and Ground Anything Anywhere at Any Granularity
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84c914e4-b932-470c-b125-8178f240c966 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Syntax-Aware Network for Handwritten Mathematical Expression Recognition
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a68251d5-6c94-4426-9053-cb39539666b5 · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e1a1a81-6f22-4224-ab42-dcc186d4266d · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Llava-grounding: Grounded visual chat with large multimodal models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 85c0c9ca-6ee4-471d-8341-845c55f471cf · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 430b7b93-879e-45fb-ab31-b9193ec284fa · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring RobuT: A systematic study of table QA robustness against human-annotated adversarial perturbations
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 75aee65c-07e5-4e9d-900c-654db90d476e · outbound
DOGR: Towards Versatile Visual Document Grounding and Referring Scale Up Composed Image Retrieval Learning via Modification Text Generation
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a83651ba-fafa-43c7-916a-56f4d0b78fbb · inbound
TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation DOGR: Towards Versatile Visual Document Grounding and Referring
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation accb5211-bd9b-4273-870a-f2ef7a501548 · inbound
DocVXQA: Context-Aware Visual Explanations for Document Question Answering DOGR: Towards Versatile Visual Document Grounding and Referring
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df04bb59-303a-4b86-b1d3-0706cfdbad84 · inbound
Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning DOGR: Towards Versatile Visual Document Grounding and Referring
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83781985-7053-40ae-a698-8e9887e53953 · inbound
DRISHTIKON: Visual Grounding at Multiple Granularities in Documents DOGR: Towards Versatile Visual Document Grounding and Referring
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 626704e1-01be-4ffa-8bbd-5cea8d2f4de8 · inbound
ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring DOGR: Towards Versatile Visual Document Grounding and Referring
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f816d463-e89f-46cf-b5f2-8a7105fdf6d0 · inbound
ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring DOGR: Towards Versatile Visual Document Grounding and Referring
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.