Pith. sign in

Paper Citation Record · LEDGER

DocVLM: Make Your VLM an Efficient Reader

As of 13 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2412.08746.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08746 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:42:14.915935Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T21:55:02.656863Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T21:59:06.400938Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact2
  • verified fuzzy27
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0d95e635-eb90-46bc-8478-43366a9a697e · outbound

This paper cites Sequence-to-sequence contrastive learning for text recogni- tion.

DocVLM: Make Your VLM an Efficient Reader Sequence-to-sequence contrastive learning for text recogni- tion

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.847032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.660000Z digest=sha256:cb387d6e0ad65cf3fd3f4e922002184ce42fb5de0ffddcecf48de4a2b5b8ea02

Observation ce335f3e-c1c6-4b33-a514-a18942a47dd9 · outbound

This paper cites Multimodal Semi-Supervised Learning for Text Recognition.

DocVLM: Make Your VLM an Efficient Reader Multimodal Semi-Supervised Learning for Text Recognition

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:42:15.286747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.664037Z digest=sha256:5df5dac2f2e40b55f040ccdaf145a1f1b4bee1ceb321b804d0a02034861d5005

Observation a83da128-c292-4c5a-82f0-a0123d3c23a5 · outbound

This paper cites Clipter: Looking at the bigger picture in scene text recogni- tion.

DocVLM: Make Your VLM an Efficient Reader Clipter: Looking at the bigger picture in scene text recogni- tion

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.832160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.667866Z digest=sha256:81de52d8109ef96d8813570fd8cc1ab91a50efd9cb9235d7661d1a1e12ea5018

Observation 97adfa2a-0880-4d59-84bf-f52aa06e0634 · outbound

This paper cites Visfocus: Prompt- guided vision encoders for ocr-free dense document under- standing.

DocVLM: Make Your VLM an Efficient Reader Visfocus: Prompt- guided vision encoders for ocr-free dense document under- standing

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.819896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.671218Z digest=sha256:455a8fae81a79dc4de0f8621e5783eb28c05d26117f33faa19a2aa2d8a8ec0d8

Observation 0dbb5cf3-6f5c-4557-bb3e-4aa26e03ea85 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

DocVLM: Make Your VLM an Efficient Reader Flamingo: a visual language model for few-shot learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.674632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.674632Z digest=sha256:b4d8b51520dc6877ec9f015ba090594d3f34b338fc4d7265fc0e5d8bde25a2e0

Observation abc368ed-db23-432e-b4c8-6a799a689b42 · outbound

This paper cites Docformer: End-to-end transformer for document understanding.

DocVLM: Make Your VLM an Efficient Reader Docformer: End-to-end transformer for document understanding

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.801855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.678011Z digest=sha256:db8d421d255b577eed1f5757aa8fbd6943ed95e0b97b402be38feee1955bbb51

Observation 33997b18-557c-491d-bfe4-53f9caf214f8 · outbound

This paper cites Docformerv2: Local features for document understanding.

DocVLM: Make Your VLM an Efficient Reader Docformerv2: Local features for document understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.777421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.681409Z digest=sha256:c473b9190170bc683d8c2186b8b83aaf7ea9b37d1e42060d72f926220b947dfc

Observation 603ac7e8-1290-4fd2-acee-b26cfdd395dd · outbound

This paper cites ScreenAI: A Vision-Language Model for UI and Infographics Understanding.

DocVLM: Make Your VLM an Efficient Reader ScreenAI: A Vision-Language Model for UI and Infographics Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.684571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.684571Z digest=sha256:e837b29d785b1d4d4301be8d2c98e38a8688294bc8ec5a9ef47c72300a9b3fdf

Observation 86291faa-31e1-4997-9aed-7da95e13ebe8 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

DocVLM: Make Your VLM an Efficient Reader Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.688335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.688335Z digest=sha256:11913ea841e30bbd3dbd37e76d285c8548a81c68d96d506fd25b665eb6786154

Observation b4053255-729c-4b5c-9baa-ecd98b1ba15f · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

DocVLM: Make Your VLM an Efficient Reader PaliGemma: A versatile 3B VLM for transfer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.691914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.691914Z digest=sha256:265c57db888748b2348226c12f981018175e47c9f320a5d7d65f979895c533ed

Observation 6e5bfb32-75f0-4d42-90f3-71b82931c738 · outbound

This paper cites Scene text visual question answering.

DocVLM: Make Your VLM an Efficient Reader Scene text visual question answering

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.761696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.695774Z digest=sha256:1938633f7252bf4a3b71596e2f96cec7cb57b0c4f47f7f50eb3a753c89e11352

Observation 680bee12-b608-4c91-9c1a-8a1c292b6423 · outbound

This paper cites Latr: Layout-aware transformer for scene-text vqa.

DocVLM: Make Your VLM an Efficient Reader Latr: Layout-aware transformer for scene-text vqa

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.739553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.699038Z digest=sha256:96d03fdeeb5b229ede5f99d95fa82d73d6425f527aa18f518117c4fadabd949d

Observation a19bad3c-e51f-4dd0-8e1b-9e85c7880300 · outbound

This paper cites OCR-IDL: OCR Annotations for Industry Document Library Dataset.

DocVLM: Make Your VLM an Efficient Reader OCR-IDL: OCR Annotations for Industry Document Library Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.702449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.702449Z digest=sha256:07cfafa2cbb5bf7e9936b207fa14a57769b7804e6c446e210330dfb15b926bae

Observation 4b3582b1-5aba-429e-9051-a6e8b13c226c · outbound

This paper cites Gram: Global reasoning for multi-page vqa.

DocVLM: Make Your VLM an Efficient Reader Gram: Global reasoning for multi-page vqa

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.706094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.706094Z digest=sha256:6813df66380491739f1f0998b7d75e071884fd3d5750520be053058a0a37789d

Observation 263231ec-21af-428e-bcf0-b8591b249bcc · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

DocVLM: Make Your VLM an Efficient Reader Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.709581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.709581Z digest=sha256:ef746df0086cd053051110184278b7169ef5b9916cd529dc6a9b478a273e64d0

Observation 70b2f9db-b7f6-457e-9d1f-2638aeb6eebb · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

DocVLM: Make Your VLM an Efficient Reader PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.713185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.713185Z digest=sha256:13dff14026a13fa1ab6312f2f2defa93a3eb721fb979196f2b373436090c7393

Observation 0db02ade-4d5c-4fc9-80c0-dd8a8cc24eb9 · outbound

This paper cites PaLI-3 Vision Language Models: Smaller, Faster, Stronger.

DocVLM: Make Your VLM an Efficient Reader PaLI-3 Vision Language Models: Smaller, Faster, Stronger

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.716518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.716518Z digest=sha256:893590f34590ca23dd4aad0aaff694ecbcc22b9bd27b58a80802d614600e3be9

Observation 3809b3e9-1c73-4ba8-b5c8-1daba241f17c · outbound

This paper cites Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy,.

DocVLM: Make Your VLM an Efficient Reader Internvl2: Better than the best—expanding performance boundaries of open-source multimodal models with the progressive scaling strategy,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.720325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.720325Z digest=sha256:75f10f92fe58b59c834a2f2eced45bc015d720e538c70a377d92fbe356c21e8d

Observation e34e2c5d-740a-430a-bb97-3b315914d230 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

DocVLM: Make Your VLM an Efficient Reader InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.723706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.723706Z digest=sha256:b6d4279696facf22b99b11bccecc218ed6edaae21f60635129056a85b40c5bb8

Observation a8a84652-7a5b-4304-b070-6d2d8702753e · outbound

This paper cites InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD.

DocVLM: Make Your VLM an Efficient Reader InternLM-XComposer2-4KHD: A Pioneering Large Vision-Language Model Handling Resolutions from 336 Pixels to 4K HD

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.727112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.727112Z digest=sha256:cc74457199bf3ec599e232e71fd8eed96e3e195817f1e4274ee00bd778ac31f5

Observation 5d08ab50-a9f8-47a5-a6ea-1497d6abc638 · outbound

This paper cites Dtrocr: Decoder-only transformer for op- tical character recognition.

DocVLM: Make Your VLM an Efficient Reader Dtrocr: Decoder-only transformer for op- tical character recognition

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.700195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.730476Z digest=sha256:8b0b22df68b21816d78eb04c31ff11c0790caaed19f2585af936766e5ce4e3e2

Observation 8e22a9ff-be1b-49c4-b305-6c5d0c8262c3 · outbound

This paper cites Towards models that can see and read.

DocVLM: Make Your VLM an Efficient Reader Towards models that can see and read

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.676824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.733829Z digest=sha256:913a128162397ff4c130d5be33b42ed662be5d2bfda4f735a2c69b5a51c62afc

Observation bb305f88-a348-4121-8b18-6bdf6dbe89d1 · outbound

This paper cites Question aware vision transformer for multimodal reasoning.

DocVLM: Make Your VLM an Efficient Reader Question aware vision transformer for multimodal reasoning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.661519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.737048Z digest=sha256:aacd583147432bd4396df447009460ef933058955b2e20dee788beccad2da547

Observation 420b070a-ff4b-4341-bc40-f9ab7456889e · outbound

This paper cites Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering.

DocVLM: Make Your VLM an Efficient Reader Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.649027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.740285Z digest=sha256:80769c507d25a94049b896fe89bb904a665b24602cbe3d819303aedd3fe38c46

Observation 74329bad-5407-48c5-8423-7176f78b40f2 · outbound

This paper cites Funsd: A dataset for form understanding in noisy scanned documents.

DocVLM: Make Your VLM an Efficient Reader Funsd: A dataset for form understanding in noisy scanned documents

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.637499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.743453Z digest=sha256:e184011d8bc22c8f8faebfbbd1db8b27884270aa27cc18a37c921657b20c0fa4

Observation c256a160-35cd-4c4f-85dc-b43b8871cd0f · outbound

This paper cites M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation.

DocVLM: Make Your VLM an Efficient Reader M3T: A New Benchmark Dataset for Multi-Modal Document-Level Machine Translation

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:42:15.161999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.748106Z digest=sha256:0e52801b6fcb0a8c8d00830112e5aedb4c1730ff9531372006c8999d7d2f5ff9

Observation 06682577-5388-4b29-b7fa-a806ecdc8dc3 · outbound

This paper cites mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding.

DocVLM: Make Your VLM an Efficient Reader mPLUG-DocOwl2: High-resolution Compressing for OCR-free Multi-page Document Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.751688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.751688Z digest=sha256:71c80b60f2de2de1356d04b823916858bc8b04114cf016602c69750095316cc8

Observation 33ef0390-6e67-4dfa-9864-6e3c4f3dac4d · outbound

This paper cites Towards unified scene text spotting based on sequence generation.

DocVLM: Make Your VLM an Efficient Reader Towards unified scene text spotting based on sequence generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.625573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.755209Z digest=sha256:45527af3c2ef00388667675a9cebbd735d54aabce57ae9c757dce8cc772ac44e

Observation e2e9d1b4-b0b2-40de-981a-93cbddac54b1 · outbound

This paper cites OCR-free Document Understanding Transformer.

DocVLM: Make Your VLM an Efficient Reader OCR-free Document Understanding Transformer

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.758608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.758608Z digest=sha256:9e0b8e4a86b64dcc026ec490064dc621b6dad517beb5232e35b8e048989baffb

Observation 7185aa3e-4c12-40d5-a219-ce5e089d5ed4 · outbound

This paper cites What matters when building vision-language models?.

DocVLM: Make Your VLM an Efficient Reader What matters when building vision-language models?

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.762029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.762029Z digest=sha256:71f320758e95261fbacd5bb5024090e81fea00bf1ab1aede9f5e304655ddeded

Observation cccefa27-388d-4548-a17e-43f726d64b16 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

DocVLM: Make Your VLM an Efficient Reader LLaVA-OneVision: Easy Visual Task Transfer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.765303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.765303Z digest=sha256:7aef03324ec3f729bb500f13339350d02d03947abfb427fe3193184a127f4f27

Observation 808b0bc4-6e05-4254-bdce-04d498a0186a · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

DocVLM: Make Your VLM an Efficient Reader Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.768805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.768805Z digest=sha256:f51394c98182e0028a55a673253b18380380308c4e6696430b914a1beba72839

Observation a0873528-fedb-49f4-88f7-8c9689af5dca · outbound

This paper cites TokenPacker: Efficient Visual Projector for Multimodal LLM.

DocVLM: Make Your VLM an Efficient Reader TokenPacker: Efficient Visual Projector for Multimodal LLM

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.772001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.772001Z digest=sha256:7476b83097276e7ad889f035390f43271bd5bb4123573ca14842c43f81167783

Observation a1cb8b79-8268-4011-8271-c5f017051d93 · outbound

This paper cites Scatter: selective con- text attentional scene text recognizer.

DocVLM: Make Your VLM an Efficient Reader Scatter: selective con- text attentional scene text recognizer

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.601849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.778183Z digest=sha256:8f3536d5105239ff77c77a9c1157cd485478fdba3ad1541bd1d080c99e77d12b

Observation 9fdad15d-55ce-4431-9fe5-4d0bf66f134d · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

DocVLM: Make Your VLM an Efficient Reader Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.781530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.781530Z digest=sha256:63e2ca29cf868f4b82a706b2bdb29ffd9ad5d0b7eae27ddb6f8c82a4ac83f887

Observation b9272f10-53a6-48da-b75c-b78946a74760 · outbound

This paper cites Visual instruction tuning.

DocVLM: Make Your VLM an Efficient Reader Visual instruction tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.784614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.784614Z digest=sha256:76e28ba95f0ce872e37948180e64e6bb133e0bd750fe1e3caa3e952fea8b1664

Observation dd5f3872-6ef9-4f77-8e34-ccbb466a6e86 · outbound

This paper cites Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024.

DocVLM: Make Your VLM an Efficient Reader Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.569569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.789406Z digest=sha256:7ecf7e98c8f41aba9341c65a44a1d0ad6a0037e3ab273c1762be2325c3855cbe

Observation 4be20eb9-54f9-4a2e-8cb7-53b5e0f58641 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

DocVLM: Make Your VLM an Efficient Reader ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.792528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.792528Z digest=sha256:753e9d03108d635decb2099f4a42cc23044af780e674d5513a3ebeed29cb6e67

Observation 83b1b8fb-4943-4c57-931b-5ba6817448e9 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

DocVLM: Make Your VLM an Efficient Reader Docvqa: A dataset for vqa on document images

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.553179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.796956Z digest=sha256:60b23fec6cee01a1b04d93475db178e5fcc62110dcd1ff8838eaedf450feb88e

Observation 49244a03-65ff-401b-8ff1-1f842a332346 · outbound

This paper cites Infographicvqa.

DocVLM: Make Your VLM an Efficient Reader Infographicvqa

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.534751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.801218Z digest=sha256:b8814e060477d8a93a27d7141f63dede21c1406e48a5823f059a4ea65115469c

Observation 19674d70-1e3f-49a8-88e9-4b1ee35b7cd7 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

DocVLM: Make Your VLM an Efficient Reader Ocr-vqa: Visual question answering by reading text in images

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.804446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.804446Z digest=sha256:40af89f4ae9682056eeef04bb9c5f5f0a21dea150f6461f48843b94e6fab55eb

Observation aec45507-a783-4669-86b1-4295e6e6e4fd · outbound

This paper cites Textadain: Paying attention to shortcut learning in text recognizers.

DocVLM: Make Your VLM an Efficient Reader Textadain: Paying attention to shortcut learning in text recognizers

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.493123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.807710Z digest=sha256:a8fdf8cb10300e025c07d85997c585dc676560f40545b5fe48e20d123365fcb7

Observation eca49a30-8850-46bd-a733-9f620a43f6f5 · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

DocVLM: Make Your VLM an Efficient Reader Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.810883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.810883Z digest=sha256:9ef59e17ee74e992e8f08f617182f62c3ee0cd6d9711bbc655f0070f19b2e9ce

Observation a1acf826-880f-4844-87cb-0403900ba6dd · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

DocVLM: Make Your VLM an Efficient Reader Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.814493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.814493Z digest=sha256:705937717c981871f1ced8b4cc87aebac3c2620b233f027f74ad762d74473308

Observation b4c32e6e-f4a5-4dfa-b566-d7031b1c6888 · outbound

This paper cites GLASS: Global to Local Attention for Scene-Text Spotting.

DocVLM: Make Your VLM an Efficient Reader GLASS: Global to Local Attention for Scene-Text Spotting

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.818997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.818997Z digest=sha256:cc9481c36660c9f44994140fe8632ecfecdafe17ce4e2da3d735efc59ee192cc

Observation 22912a0c-9d61-4879-b8fe-13698debcdd3 · outbound

This paper cites Textcaps: a dataset for image caption- ing with reading comprehension.

DocVLM: Make Your VLM an Efficient Reader Textcaps: a dataset for image caption- ing with reading comprehension

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.822578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.822578Z digest=sha256:b425a47452e41e91079f8a2e1884f6de897111c8a1cb3218b3078e50c2ddca2e

Observation 3d30a0a9-8f37-44cb-99c9-20d469c48152 · outbound

This paper cites Towards vqa models that can read.

DocVLM: Make Your VLM an Efficient Reader Towards vqa models that can read

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.438395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.826725Z digest=sha256:42859cd03bb1a39dd94b0afb86a42a65c8372290e4c07c92770e66e9bca40759

Observation e2a6e898-08b5-4869-b4ea-79fc57b94d29 · outbound

This paper cites Instructdoc: A dataset for zero-shot gener- alization of visual document understanding with instructions.

DocVLM: Make Your VLM an Efficient Reader Instructdoc: A dataset for zero-shot gener- alization of visual document understanding with instructions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.421608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.832432Z digest=sha256:bfcaa8fbee0d181934a702ac17ce7d32563fa8373abe0d74c8c7e88f587dfea3

Observation 4fc6056f-ba81-4d87-9163-bd3fbb4aa3c4 · outbound

This paper cites Hi- erarchical multimodal transformers for multipage docvqa.

DocVLM: Make Your VLM an Efficient Reader Hi- erarchical multimodal transformers for multipage docvqa

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.391807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.836449Z digest=sha256:e3660d69a44d10e3fa42797e7f8c6dfc9357fd2f31154d5591267d713ed0f34c

Observation f2f38ca2-8ee0-429e-a95b-d9d29bfbb79a · outbound

This paper cites Document understanding dataset and evaluation (dude).

DocVLM: Make Your VLM an Efficient Reader Document understanding dataset and evaluation (dude)

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.370799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.850835Z digest=sha256:caa6ca936b74668f7ec9119f65430977ad76267fc9dabe0a3c950bbc064c5ca2

Observation 8a1ab35e-680a-4395-8b0d-539f5af5fb83 · outbound

This paper cites DocLLM: A layout-aware generative language model for multimodal document understanding.

DocVLM: Make Your VLM an Efficient Reader DocLLM: A layout-aware generative language model for multimodal document understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.858346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.858346Z digest=sha256:42723c9ef3106fcde4c12e2cbe2f7c574fbb9199721480e4b63b8752a50d706c

Observation c6e054a9-6d96-4299-8b1f-da5b50f2d127 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

DocVLM: Make Your VLM an Efficient Reader Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.862395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.862395Z digest=sha256:982e65a6434363a68fedd930af845a20088dcd55610a49fb5f02db54c0397701

Observation c96d11ff-42d2-4c2a-9ca8-414d0cfd81f1 · outbound

This paper cites Layout and Task Aware Instruction Prompt for Zero-shot Document Image Question Answering.

DocVLM: Make Your VLM an Efficient Reader Layout and Task Aware Instruction Prompt for Zero-shot Document Image Question Answering

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.867758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.867758Z digest=sha256:89dce53cf0d03b57f193b255dd9d89f7cffd6e7fa9ae48e913f9a495a92c220b

Observation b02716ba-9456-4415-9e20-88f23408f67f · outbound

This paper cites Layoutlm: Pre-training of text and layout for document image understanding.

DocVLM: Make Your VLM an Efficient Reader Layoutlm: Pre-training of text and layout for document image understanding

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.359252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.873685Z digest=sha256:886f83f05539f4a85097ff5895bdf8cb2491f59da99b1b056ac18a0e1e44b332

Observation 9f687888-e6cb-4ea7-8640-113138fd9b41 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

DocVLM: Make Your VLM an Efficient Reader UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.877803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.877803Z digest=sha256:8352bc985d2ed2f2da30927eef0505e76dd5870ba1f51fe09c64da80d804d93f

Observation c9cca9cb-c68d-4cc0-8bb8-54592e6afc9e · outbound

This paper cites Dptext-detr: Towards better scene text detection with dynamic points in transformer.

DocVLM: Make Your VLM an Efficient Reader Dptext-detr: Towards better scene text detection with dynamic points in transformer

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.347097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.881568Z digest=sha256:2b3e8eea8b32591f561bfc7c862815c1d0e74bdc372e15b03e41ee7a5345241a

Observation a6f9ff17-c632-449a-b020-fb0fb9f84129 · outbound

This paper cites Deepsolo: Let transformer decoder with explicit points solo for text spot- ting.

DocVLM: Make Your VLM an Efficient Reader Deepsolo: Let transformer decoder with explicit points solo for text spot- ting

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.331773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.888590Z digest=sha256:3cff5ad588d549a7aef2d82eb668f5ff16d805298dfffb9d3b71ecfdc9eb9b51

Observation 4d194bfd-b861-4324-b4ef-628cc0342fc5 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

DocVLM: Make Your VLM an Efficient Reader mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.892017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.892017Z digest=sha256:72e8780fce4d79a791dc8c71ad965b728273706639d7ddf30aff0e0298c86e32

Observation 381c7c14-7bc2-468b-adad-b45847ef488a · outbound

This paper cites InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output.

DocVLM: Make Your VLM an Efficient Reader InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.897792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.897792Z digest=sha256:4f6039526ba82072d1468d5293fc4a4b829241e2b582c2151b9d762100f85990

Observation f3afa750-7b12-474b-ae8e-56f25805e8af · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

DocVLM: Make Your VLM an Efficient Reader MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T17:42:14.901281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:42:14.901281Z digest=sha256:c35dc7669af999659a5f40e3b088b30176e5ef2f925ead2138861b16b4af0b59

Observation 23c2e6cd-2f15-4de5-b2ee-35443e785d9d · outbound

This paper cites Towards complex doc- ument understanding by discrete reasoning.

DocVLM: Make Your VLM an Efficient Reader Towards complex doc- ument understanding by discrete reasoning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.318670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.905475Z digest=sha256:9f9be6cf5d876e9e24e1f6c51144adbba723d14a5e47055d7b8649ce3337f3ec

Observation c79542d8-7131-404a-813b-9090a688b5c9 · outbound

This paper cites The encoder is initial- ized with pretrained weights from DocFormerV2, which was pretrained on the Industry Document Library (IDL) dataset [13].

DocVLM: Make Your VLM an Efficient Reader The encoder is initial- ized with pretrained weights from DocFormerV2, which was pretrained on the Industry Document Library (IDL) dataset [13]

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:42:15.302121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T17:42:14.915935Z digest=sha256:c509b1aff19c6050bcf1da2fc4f38cdb59791f0f0485618be2f36986054564f9

Pith citing papers

Observation 83f55883-6cbf-4800-a317-78e8c3ccad98 · inbound

Operationalizing Document AI: A Microservice Architecture for OCR and LLM Pipelines in Production cites this paper.

Operationalizing Document AI: A Microservice Architecture for OCR and LLM Pipelines in Production DocVLM: Make Your VLM an Efficient Reader

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:59:06.404665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T21:55:02.656863Z digest=sha256:8cdd665ff6ed72d199987cda96706f3e837fcdbcf92aff7f49873cced3ef37e3