Pith. sign in

Paper Citation Record · LEDGER

DOGR: Towards Versatile Visual Document Grounding and Referring

As of 18 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 6 inbound Pith citation observations for arXiv:2411.17125.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17125 v3

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:34:31.918078Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:09:11.055077Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T23:35:07.388169Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26bfbaeb-8f0a-4d97-81af-e54ea2c53471 · outbound

This paper cites Qwen2.5-VL Technical Report.

DOGR: Towards Versatile Visual Document Grounding and Referring Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.648211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.648211Z digest=sha256:fae68d924c2bad6249e9ffd096d8f43ab9867fd77c065601e970b03b34497ade

Observation 24a4114f-9514-43c0-99d1-d27c2a3f1906 · outbound

This paper cites Jawahar, Ernest Valveny, and Dimos- thenis Karatzas.

DOGR: Towards Versatile Visual Document Grounding and Referring Jawahar, Ernest Valveny, and Dimos- thenis Karatzas

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.916172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.653462Z digest=sha256:9722b77874ed99fa4b5bfde5ccde676cea774d39393da2d804d2c4209d0949a5

Observation 3bdda016-7e96-44d9-82cc-85788fd7cd7a · outbound

This paper cites Textocr-gpt4v.

DOGR: Towards Versatile Visual Document Grounding and Referring Textocr-gpt4v

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.902951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.657301Z digest=sha256:ece58ab46e6474f372532b927bbcb619b4f22556da93f6d505ab472939a2837d

Observation 17f76eeb-cb04-45da-ac8d-ea8681731f84 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

DOGR: Towards Versatile Visual Document Grounding and Referring Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.661116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.661116Z digest=sha256:b6bb3e6b3905a0e6fa3eb308154a5f4ce8762e9b5eb0a72ce3cb6fb47063affe

Observation 6bd12d6c-c3d4-4aa0-aeb2-f7bcec147bfb · outbound

This paper cites Tabfact : A large-scale dataset for table-based fact verification.

DOGR: Towards Versatile Visual Document Grounding and Referring Tabfact : A large-scale dataset for table-based fact verification

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.890189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.665898Z digest=sha256:4c61f3aadce68c86357ada2042ab2a342bcf82c1e676292523c6da6413cf61b3

Observation 6a3282b8-7a6b-4543-9a1f-523f3ca78d66 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

DOGR: Towards Versatile Visual Document Grounding and Referring InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.669626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.669626Z digest=sha256:084592f23a0bacb3e9908ec354b9f863b979524ddd2da185bde36a51892c15fd

Observation 11801deb-4f35-47ee-946e-a8a415b08cbe · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

DOGR: Towards Versatile Visual Document Grounding and Referring Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.673783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.673783Z digest=sha256:624d2db7b41e623eba542db4fd300c7a9365676ce97ec4d60f71b75a8c0d100e

Observation 1ae2b981-6d89-453b-96ca-01ddd952001c · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

DOGR: Towards Versatile Visual Document Grounding and Referring How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.678432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.678432Z digest=sha256:8f4269c2fac76e65c6895e3339efb15adcbad08af3b5a3600fa7987f602be2ff

Observation b5be5ab8-9bc1-44a1-add0-7135206b6aa0 · outbound

This paper cites Hitab: A hierarchical table dataset for question an- swering and natural language generation.

DOGR: Towards Versatile Visual Document Grounding and Referring Hitab: A hierarchical table dataset for question an- swering and natural language generation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.877411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.682552Z digest=sha256:75effe7c97c5484e2f522d3772ae79c50d80c92dac8bdbc5f072bf702a6ccc64

Observation 0833509a-d22a-43bd-b5b2-dc5b3badfb5a · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

DOGR: Towards Versatile Visual Document Grounding and Referring Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.686379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.686379Z digest=sha256:73f374840ca667702b643dc344b66ae90629fadf2d9839afa5946b47047f04fa

Observation 86c18ed8-c0ac-4dd9-8bae-979d92f363c7 · outbound

This paper cites Pymupdf: Python bindings for mupdf.

DOGR: Towards Versatile Visual Document Grounding and Referring Pymupdf: Python bindings for mupdf

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.864954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.690243Z digest=sha256:59a11d3a845e903c801052853c681e5fae8fcfd0503cbcb55ee21cc2b9234f04

Observation 02369ccb-e3bd-494a-a328-777d4ed14c00 · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

DOGR: Towards Versatile Visual Document Grounding and Referring InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.694114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.694114Z digest=sha256:72b8244d4931ce200278bdc5d1b3118cba1bda402f9837321f131aaa1f6ca41a

Observation ff194cb4-9bee-4e45-ab8c-d9081b52d5a2 · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

DOGR: Towards Versatile Visual Document Grounding and Referring mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.702415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.702415Z digest=sha256:eff872e0891f606332d58bab8c9f7bae842f5d645a2e4a2aa988f1154fcdae48

Observation 3c2d8526-7b32-4d4b-ab24-bd8dbacfbb5b · outbound

This paper cites mplug- docowl2: High-resolution compressing for ocr-free multi- page document understanding, 2024.

DOGR: Towards Versatile Visual Document Grounding and Referring mplug- docowl2: High-resolution compressing for ocr-free multi- page document understanding, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.851190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.706461Z digest=sha256:68f8384bf5a4958957ee0479071c1d556b4abcdfad51a4adc5c90b6fbc84fce3

Observation c4b5a9a1-23d7-4426-9464-4172398ab6e7 · outbound

This paper cites GPT-4o System Card.

DOGR: Towards Versatile Visual Document Grounding and Referring GPT-4o System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.710462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.710462Z digest=sha256:0fc56fde1f5c3a5f13caba7286765ed87d20f5c942198fd56f2917503add7ff5

Observation 685b00ea-5278-4467-b538-07ddaa96579e · outbound

This paper cites Dvqa: Understanding data visualizations via ques- tion answering.

DOGR: Towards Versatile Visual Document Grounding and Referring Dvqa: Understanding data visualizations via ques- tion answering

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.838579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.714509Z digest=sha256:bc2249cbc1341a8654096c47b63c7513a1b588b5c9e4a8b339ae8f17592483f8

Observation bf19eda8-5938-4e02-9bbe-f4aa97156073 · outbound

This paper cites Fig- ureqa: An annotated figure dataset for visual reasoning,.

DOGR: Towards Versatile Visual Document Grounding and Referring Fig- ureqa: An annotated figure dataset for visual reasoning,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.718388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.718388Z digest=sha256:c771d807f5431fdfb25d71e6a7d8255ffefdb94c6375d4a74851ae3e3e2aef1d

Observation abc404f9-3788-4835-88f4-690e59522ef0 · outbound

This paper cites A diagram is worth a dozen images.

DOGR: Towards Versatile Visual Document Grounding and Referring A diagram is worth a dozen images

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.818590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.722309Z digest=sha256:fdeeb69612c19445b33f085811f1ae066753223f5f16382cb4cd7129579842d9

Observation 3201579a-ddad-4b98-bd1a-cf897857c69a · outbound

This paper cites Are you smarter than a sixth grader? textbook question answer- ing for multimodal machine comprehension.

DOGR: Towards Versatile Visual Document Grounding and Referring Are you smarter than a sixth grader? textbook question answer- ing for multimodal machine comprehension

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.806395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.726598Z digest=sha256:faa18f777c381b0a9fbf38e37980e0906f66e35f989aca4a8922d60323a83802

Observation 8254ade0-206f-4b5e-bed7-4b9587714bb1 · outbound

This paper cites Ocr-free document understanding transformer.

DOGR: Towards Versatile Visual Document Grounding and Referring Ocr-free document understanding transformer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.793430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.730781Z digest=sha256:946db9ae45461a8cfd082628cd50abd816cf8353b3518e990ad3c4a79dbac2ae

Observation bd41f984-f577-42d4-a053-51f817060538 · outbound

This paper cites TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation.

DOGR: Towards Versatile Visual Document Grounding and Referring TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.734719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.734719Z digest=sha256:f4f51d9d668415a606fb07e458fb6224928d401aee887fb04944e845f0497c5e

Observation 233f6a0a-37ed-4caf-b440-9f7b4dbe29bb · outbound

This paper cites Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want,.

DOGR: Towards Versatile Visual Document Grounding and Referring Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.781648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.738941Z digest=sha256:a03633e2c29bc1e702558b4212cf72b4797550f5ba7d7dd97159ab10f5edd07d

Observation f8c0dc91-a888-4c5e-8823-ec4ddd7802b6 · outbound

This paper cites Focus Anywhere for Fine-grained Multi-page Document Understanding.

DOGR: Towards Versatile Visual Document Grounding and Referring Focus Anywhere for Fine-grained Multi-page Document Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.746755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.746755Z digest=sha256:641461bb7b5ef74594da1b5f98e76c3fc0bf7da60590155e110e329ae2c151aa

Observation 82070d9c-33c0-4d3d-a63d-994a987b1975 · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

DOGR: Towards Versatile Visual Document Grounding and Referring Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.750312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.750312Z digest=sha256:2a7d1e0558c57ce856f1e77cba182766bce61f92b753d180f833bd24db72da23

Observation b86853e7-b33e-4e6d-a01f-1ae64ff2d84f · outbound

This paper cites MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning.

DOGR: Towards Versatile Visual Document Grounding and Referring MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.754314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.754314Z digest=sha256:bbadba9984b0834c4517cd4fcebd105632ca0272e9c0312e68356ef72720da7d

Observation 3a49c185-310c-466b-9d49-ddf16efa2cb9 · outbound

This paper cites Visual instruction tuning, 2023.

DOGR: Towards Versatile Visual Document Grounding and Referring Visual instruction tuning, 2023

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.758827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.758827Z digest=sha256:f5c13edd3947b9a0838702f1af3d1997e6ea6c9960e218d0b900fadcd9c66639

Observation cf4bc3c0-0fc8-4a09-808e-3e2358f57ebe · outbound

This paper cites Textmonkey: An ocr-free large multimodal model for understanding document, 2024.

DOGR: Towards Versatile Visual Document Grounding and Referring Textmonkey: An ocr-free large multimodal model for understanding document, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.762840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.763097Z digest=sha256:5c204e6707a185ba7b0d3c8fd09f1d6785820a65f513fee1349b1883da024ba2

Observation 55824c0c-9478-463f-8219-ebe694609b52 · outbound

This paper cites KOSMOS-2.5: A Multimodal Literate Model.

DOGR: Towards Versatile Visual Document Grounding and Referring KOSMOS-2.5: A Multimodal Literate Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.768116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.768116Z digest=sha256:e8234124f36c2b81c473764501faabffc76479fe969f305a7118fa490d7f2ce6

Observation b46eb2d2-7fa6-4847-8a65-a189fd7ff92f · outbound

This paper cites The iam-database: an english sentence database for offline handwriting recognition.

DOGR: Towards Versatile Visual Document Grounding and Referring The iam-database: an english sentence database for offline handwriting recognition

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.750409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.772448Z digest=sha256:1bf54bb655d419fb4abb15d050e7adc851a9c4b2dd0626d12e9504bb957d9769

Observation b971a719-8ed5-4cd5-801f-d4ba6aa2f63c · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

DOGR: Towards Versatile Visual Document Grounding and Referring ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.737825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.776555Z digest=sha256:abc3475cdaa99fdfdcfb18ee9e132cc5f4add06983ba10db4df1cefa5e128144

Observation cfb35cac-fa5c-485d-8c02-57a3c615389e · outbound

This paper cites Chartqa: A benchmark for question an- swering about charts with visual and logical reasoning.

DOGR: Towards Versatile Visual Document Grounding and Referring Chartqa: A benchmark for question an- swering about charts with visual and logical reasoning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.723847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.780990Z digest=sha256:355fe908c3b831c188d50e9d49b158fea17a08a46192171c24b002cbd3181e0d

Observation 4d854a3b-a335-4c5b-b400-fb0d838884a2 · outbound

This paper cites an unresolved cited work.

DOGR: Towards Versatile Visual Document Grounding and Referring Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:34:32.710777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.784819Z digest=sha256:0ecff2e07aa93bdea3422f3ee1a2b66cd43631da3d75d71497dc6c1493a56036

Observation fc59c0a9-032d-41cc-917b-9dcc68d1af8a · outbound

This paper cites an unresolved cited work.

DOGR: Towards Versatile Visual Document Grounding and Referring Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-12T12:34:32.698912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.789514Z digest=sha256:123d5420a97fcd27b9582f35c2a716f15c6f2f31b068d641f86425e1b5d171f2

Observation ff135251-6a65-4408-b668-52e9ca751260 · outbound

This paper cites Mishra, K.

DOGR: Towards Versatile Visual Document Grounding and Referring Mishra, K

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.686897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.793645Z digest=sha256:e6f23fc72ee43cb24ba7c275ff2a556cb7108c7aae527ca7edbbdf3f5166a26f

Observation f8b951c5-e9b6-4b9e-ad33-16912428e8d7 · outbound

This paper cites Chart-to-text: Generat- ing natural language descriptions for charts by adapting the transformer model, 2020.

DOGR: Towards Versatile Visual Document Grounding and Referring Chart-to-text: Generat- ing natural language descriptions for charts by adapting the transformer model, 2020

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.674676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.798015Z digest=sha256:74296a01f6e66ed0f3377c6765be1bf973200fad75100f37c3f83ba53e556870

Observation 73c958f6-3054-4c68-aedd-73337ca00cd2 · outbound

This paper cites Compositional semantic parsing on semi-structured tables.

DOGR: Towards Versatile Visual Document Grounding and Referring Compositional semantic parsing on semi-structured tables

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.662272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.801879Z digest=sha256:07de2c65d42533a90ce47dce9a3be598898d3c8b5c4ef0905878903772985836

Observation bcb6e8df-9dc2-4b8c-85fc-5b4f0032ac27 · outbound

This paper cites Kosmos-2: Ground- ing multimodal large language models to the world.

DOGR: Towards Versatile Visual Document Grounding and Referring Kosmos-2: Ground- ing multimodal large language models to the world

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.805555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.805555Z digest=sha256:3ffe56d4b9c206bbbd56dfe67bf52848c2dc1c69bb496f1f65626ef4a983668d

Observation c4006fc3-7d27-4cd1-94da-e727d5633d65 · outbound

This paper cites Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S.

DOGR: Towards Versatile Visual Document Grounding and Referring Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.641077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.809155Z digest=sha256:d9daa39cc3afc6900c1478068d4037c64872ef24fb7500a5fb23f2aa076cafc4

Observation 139bd0b3-a96d-496f-9f22-13e542788837 · outbound

This paper cites Textcaps: a dataset for image caption- ingwith reading comprehension.

DOGR: Towards Versatile Visual Document Grounding and Referring Textcaps: a dataset for image caption- ingwith reading comprehension

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.627247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.812360Z digest=sha256:4086fa292964a79a80e4102a2854cadd47179498a8eed2b14f9d38e2ddd33784

Observation 6aee4a9f-96da-4d56-91db-0d9ac003b27e · outbound

This paper cites Towards vqa models that can read.

DOGR: Towards Versatile Visual Document Grounding and Referring Towards vqa models that can read

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.614643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.815734Z digest=sha256:ee56910da5456410d39a4eb3a90156ebea0aada3c9fe10bfd277e0369b0bd717

Observation dd376bc9-496b-41e7-9f7c-a4a872091e98 · outbound

This paper cites Kleister: Key in- formation extraction datasets involving long documents with complex layouts.

DOGR: Towards Versatile Visual Document Grounding and Referring Kleister: Key in- formation extraction datasets involving long documents with complex layouts

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.468596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.819457Z digest=sha256:bf039a618d21797eaf0912eb481ee13a342437565c22274eae21cd51065d28c7

Observation 53f13a80-5f57-42a4-b9d7-25d681224026 · outbound

This paper cites Deepform: Understand structured docu- ments at scale.

DOGR: Towards Versatile Visual Document Grounding and Referring Deepform: Understand structured docu- ments at scale

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.454889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.823183Z digest=sha256:68b0146fc976b9a9a24beb88fe74f3178bfacd56c48ea768184fd3404c3e4c91

Observation b6e509b7-96b3-4561-9974-20b284045ab9 · outbound

This paper cites Vi- sualmrc: Machine reading comprehension on document im- ages.

DOGR: Towards Versatile Visual Document Grounding and Referring Vi- sualmrc: Machine reading comprehension on document im- ages

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.436897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.827900Z digest=sha256:ddb64137579bd630add00054620ba5e516af0aa2fccfc71a7c1dba17327f9d9f

Observation 16a1f5ae-9a57-492d-ab18-c1ea2f67bf5d · outbound

This paper cites Tang, Angie Boggust, and Arvind Satyanarayan.

DOGR: Towards Versatile Visual Document Grounding and Referring Tang, Angie Boggust, and Arvind Satyanarayan

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.831664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.831664Z digest=sha256:4fbe0b8ff8727ee1eb8f09c60ba92b1a4f7c3643e43692b9704de3e85bd684ee

Observation 4a88d28f-3ed9-43c3-9fab-4cf55b9231c2 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

DOGR: Towards Versatile Visual Document Grounding and Referring Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.836067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.836067Z digest=sha256:a1d74f58c6ddc4f53f16b1c97f870c2e76b3651fdd05fd756eb04b93cd18c64b

Observation c97f8457-39f0-4a02-b8e9-5a0f22fd0dcb · outbound

This paper cites Cc-main-2021-31-pdf- untruncated.

DOGR: Towards Versatile Visual Document Grounding and Referring Cc-main-2021-31-pdf- untruncated

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.412470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.840254Z digest=sha256:85be03a0604281f1b338cbed3817e7e19289383bd26077bd92f89b7a8db5853d

Observation 73d5a102-9789-401d-ad47-2de3a29363ea · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

DOGR: Towards Versatile Visual Document Grounding and Referring Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.844044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.844044Z digest=sha256:da1d310b0d4aff52a8e0855e3f2ed03cccf862da9a5d732bf813b5b9513c9225

Observation 002e889c-ada0-4abe-8e71-2fd052a4fbfe · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

DOGR: Towards Versatile Visual Document Grounding and Referring Lawrence Zitnick, and Devi Parikh

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.399990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.848176Z digest=sha256:56e7a293b86e987404852ec98bfa472c6a2e55cf42f008d970b382229d302379

Observation 833c0e39-612c-45fd-992d-e3ab4ab6ccd2 · outbound

This paper cites Screen2words: Automatic mobile ui summarization with multimodal learning, 2021.

DOGR: Towards Versatile Visual Document Grounding and Referring Screen2words: Automatic mobile ui summarization with multimodal learning, 2021

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.388039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.851720Z digest=sha256:a34ee57ef0156b6dae365b66f54dcadfd789b0bb83f13528dc49d625015e65cc

Observation a6ae5d35-1d18-43a3-8b4e-ec6e9f6e336c · outbound

This paper cites Mineru: An open-source solution for precise document content extrac- tion, 2024.

DOGR: Towards Versatile Visual Document Grounding and Referring Mineru: An open-source solution for precise document content extrac- tion, 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.375706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.855418Z digest=sha256:407538a728b6d802019a559cd4b45a790787d63ddcf6d4c6371ffd4be17129d6

Observation fb01f736-d8c8-492b-8506-5086074fa2e6 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

DOGR: Towards Versatile Visual Document Grounding and Referring Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.859533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.859533Z digest=sha256:268da0bdf837d42c4591fbec570af4184e8b7fb7c6bbf9b8b55d79882f82036a

Observation 03780143-2616-445c-ba74-cfe102edf024 · outbound

This paper cites Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models.

DOGR: Towards Versatile Visual Document Grounding and Referring Vary: Scaling up the Vision Vocabulary for Large Vision-Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.863650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.863650Z digest=sha256:9fd514713d0badcb725c05b03e4b50fbf6bc98002d5b73d5b74149152be8f86a

Observation 91f13e26-c975-4900-a338-4ae045535b86 · outbound

This paper cites wendlerc/renderedtext, 2023.

DOGR: Towards Versatile Visual Document Grounding and Referring wendlerc/renderedtext, 2023

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.362511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.868240Z digest=sha256:0c2c67f771985c20d68d1ed16365a6c40b5e2401b37dded0b24ffb79bb39a873

Observation 39990f8d-4f61-4327-acb1-f0f1310396da · outbound

This paper cites Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing.

DOGR: Towards Versatile Visual Document Grounding and Referring Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.873021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.873021Z digest=sha256:dd4e02de1f599a4b04495d2bc7882c916609aee4de9a373f2973e2e9298f6aa8

Observation 397fcff5-cc26-4214-9c8a-d9b81bec51da · outbound

This paper cites Canvasvae: Learning to generate vector graphic documents.

DOGR: Towards Versatile Visual Document Grounding and Referring Canvasvae: Learning to generate vector graphic documents

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.877150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.877150Z digest=sha256:1b260e170808fb023ebdcfd135373a19efd52edb7f14ca27d6447a60d3ffa27e

Observation 2f46af81-460d-4109-9a0e-016a452736b2 · outbound

This paper cites Qwen2 Technical Report.

DOGR: Towards Versatile Visual Document Grounding and Referring Qwen2 Technical Report

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.881839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.881839Z digest=sha256:5f82bb199c8772a9293aa062506a6be3f426e46c3bb155a34b447138f1d9660d

Observation e472e1d7-8b69-4fb1-8f6c-ecbb9889acab · outbound

This paper cites Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model.

DOGR: Towards Versatile Visual Document Grounding and Referring Ureader: Universal ocr-free visually-situated language understanding with multimodal large language model

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.340509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.886664Z digest=sha256:d4e57bc7ca72298e0620d1e354d09bc35e84537e3f00af1bf6748799c4354309

Observation 1852f1e3-5581-4968-921f-5172e00bd02c · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

DOGR: Towards Versatile Visual Document Grounding and Referring Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.890922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.890922Z digest=sha256:b8e86ac767332cd79d0c04f35cce7667e6871d04d968bbdebb4345d088511ad0

Observation 84c914e4-b932-470c-b125-8178f240c966 · outbound

This paper cites Syntax-Aware Network for Handwritten Mathematical Expression Recognition.

DOGR: Towards Versatile Visual Document Grounding and Referring Syntax-Aware Network for Handwritten Mathematical Expression Recognition

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.895427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.895427Z digest=sha256:cf86c6ebd399a217555ada3c5d59b645f00781b35832e09ea71acce559355d1f

Observation a68251d5-6c94-4426-9053-cb39539666b5 · outbound

This paper cites MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning.

DOGR: Towards Versatile Visual Document Grounding and Referring MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.900255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.900255Z digest=sha256:546b2c1eea3e3fecf881cf66db9ce0665a2a0911536ef928cce547d896b8f847

Observation 9e1a1a81-6f22-4224-ab42-dcc186d4266d · outbound

This paper cites Llava-grounding: Grounded visual chat with large multimodal models.

DOGR: Towards Versatile Visual Document Grounding and Referring Llava-grounding: Grounded visual chat with large multimodal models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.326936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.904576Z digest=sha256:80625deb93790c7122d4e62f82e1955d5d436af7757142067b06b2963f12f4ad

Observation 85c0c9ca-6ee4-471d-8341-845c55f471cf · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

DOGR: Towards Versatile Visual Document Grounding and Referring InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.909206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.909206Z digest=sha256:c5dc0067652afcfe3d93261992fbaf6d8e02dc1bf8ea3d16f639515246daf4ce

Observation 430b7b93-879e-45fb-ab31-b9193ec284fa · outbound

This paper cites RobuT: A systematic study of table QA robustness against human-annotated adversarial perturbations.

DOGR: Towards Versatile Visual Document Grounding and Referring RobuT: A systematic study of table QA robustness against human-annotated adversarial perturbations

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:34:32.314568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T12:34:31.913697Z digest=sha256:2ad916762ea3d2731503cfe880f85c89c99946531c092b9cbcb8c1a3d2a3a374

Observation 75aee65c-07e5-4e9d-900c-654db90d476e · outbound

This paper cites Scale Up Composed Image Retrieval Learning via Modification Text Generation.

DOGR: Towards Versatile Visual Document Grounding and Referring Scale Up Composed Image Retrieval Learning via Modification Text Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T12:34:31.918078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:34:31.918078Z digest=sha256:de677c5679311071d88f7780cc8ab389851c0db40028cb78cfd7adff7eb4660e

Pith citing papers

Observation a83651ba-fafa-43c7-916a-56f4d0b78fbb · inbound

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation cites this paper.

TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation DOGR: Towards Versatile Visual Document Grounding and Referring

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T23:09:11.055077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:09:11.055077Z digest=sha256:7ff6275132378b63d6a525a9b3d042e435ba25e56a00f37e91506574a48b8e74

Observation accb5211-bd9b-4273-870a-f2ef7a501548 · inbound

DocVXQA: Context-Aware Visual Explanations for Document Question Answering cites this paper.

DocVXQA: Context-Aware Visual Explanations for Document Question Answering DOGR: Towards Versatile Visual Document Grounding and Referring

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T22:20:55.593869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:20:55.593869Z digest=sha256:a04b49b261259495482dd0f485b1c2e1a8caf1ee23248247bc2e13b4cb0a575a

Observation df04bb59-303a-4b86-b1d3-0706cfdbad84 · inbound

Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning cites this paper.

Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning DOGR: Towards Versatile Visual Document Grounding and Referring

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:40.427096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:40.427096Z digest=sha256:2832af40b114144fe4f050589238fa190c3a84e0275a594b5313630b6bbbafd2

Observation 83781985-7053-40ae-a698-8e9887e53953 · inbound

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents cites this paper.

DRISHTIKON: Visual Grounding at Multiple Granularities in Documents DOGR: Towards Versatile Visual Document Grounding and Referring

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:34:38.340134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:34:38.340134Z digest=sha256:17ca447548732a7299f69e5a1e5ceccd6c2c96a55194f2465402f2a5499dc018

Observation 626704e1-01be-4ffa-8bbd-5cea8d2f4de8 · inbound

ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring cites this paper.

ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring DOGR: Towards Versatile Visual Document Grounding and Referring

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:20:56.645771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-11T01:48:09.517875Z digest=sha256:af766912dbd486f05f0e2f80f78b33a916e56a079f3b62460abaa8b23fc3c44f

Observation f816d463-e89f-46cf-b5f2-8a7105fdf6d0 · inbound

ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring cites this paper.

ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse referring clues and Multi-Target Referring DOGR: Towards Versatile Visual Document Grounding and Referring

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:35:07.389489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T23:31:19.715414Z digest=sha256:d9cb844e02a7c1ea9812d276d26eba725056508813f4c20b1368d2f272a2a701