Pith. sign in

Paper Citation Record · LEDGER

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization

As of 15 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2507.09531.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09531 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:59:07.601054Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29dfd4c0-64c6-4a2c-acdd-19fef097dacb · outbound

This paper cites Docformer: End-to-end transformer for document understanding.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Docformer: End-to-end transformer for document understanding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:15.872435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:02.282868Z digest=sha256:5d8d91fa8674103d6f1061b0d19a96c6995cdb21fc6abcc387f8d3a1048f0a11

Observation 5508c845-f007-4386-b509-bdb252ee4268 · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Qwen-vl: A versatile vision-language model for un- derstanding, localization

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:15.717043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:02.385870Z digest=sha256:75417065c48d3708eec2c1d5e137b78124267f410ebea8cdc0ce1417dcfd4783

Observation 39ece4dc-3400-49d9-a18a-1fba010b1dc6 · outbound

This paper cites Due: End-to-end document understand- ing benchmark.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Due: End-to-end document understand- ing benchmark

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:15.610275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:02.495887Z digest=sha256:7889c38180327d5853da70552eaa73cc04fda2bfecb23281eff4cb88d0ac72b8

Observation e0ffee70-b30d-4c9a-a799-f5fae76f8f5d · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:02.588039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:02.588039Z digest=sha256:8a412880004ce3fa8533edf072c8362ddde3a434c30d01bcf37eb128a04ed055

Observation be0e5316-1b31-4aa0-94eb-a6704667ce56 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Gonzalez, Ion Stoica, and Eric P

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:15.523868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:02.710552Z digest=sha256:d43dcf06f1e912309f0f57243acec89df7d53a23819822c7475a61b7356ed4da

Observation ad2367e2-b82d-45f9-8574-151d13ee6d98 · outbound

This paper cites Instructblip: towards general- purpose vision-language models with instruction tuning.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Instructblip: towards general- purpose vision-language models with instruction tuning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:15.377660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:02.852460Z digest=sha256:c969148a152e2359e027956b057aae2365c2759b3e2757a861bc07d5299414ff

Observation 2562b1e0-86bb-4dd2-b276-104e062ac7ab · outbound

This paper cites Docpedia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Docpedia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:15.224822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:03.002811Z digest=sha256:3e6ce68ab134fe19b43a270b2acb1f8fd153fae550f0ae2238ef0cd1f48ae373

Observation 72ea372f-a773-49ab-b868-f35ea1b0ddd1 · outbound

This paper cites Unidoc: Unified pretraining framework for document understanding.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Unidoc: Unified pretraining framework for document understanding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:15.072082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:03.160371Z digest=sha256:9d200b3d2fdc490ff41e525edc381726889ecc33f4c00db0d8351feaeb8800ab

Observation e1a838da-28c7-4965-8f29-5a4561e6547e · outbound

This paper cites Deep residual learning for image recognition.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Deep residual learning for image recognition

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:14.952677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:03.291969Z digest=sha256:ebaaa75d36b6312fcb0351f513fe1fac0679bb4b74cfdfdd505a3389e0cebd4d

Observation b584473d-7079-4be1-ace2-70cadbd24256 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Gaussian Error Linear Units (GELUs)

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:03.420452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:03.420452Z digest=sha256:48aff04bf41573314ab7a9eb734b8634043455700084deefd3ff3ec57c21a680

Observation 1fec0b0f-ce6d-4e54-ad0f-301f5e9fcad5 · outbound

This paper cites Cogagent: A visual language model for gui agents.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Cogagent: A visual language model for gui agents

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:14.780291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:03.568961Z digest=sha256:49bdf090adf4631647d656ab0758e3d571c0573325e1efbd1f51e566e981e9df

Observation 33ee729d-cb83-4a84-8245-2eb18fe80198 · outbound

This paper cites SciCap: Generating captions for scientific figures.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization SciCap: Generating captions for scientific figures

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:14.636710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:03.706824Z digest=sha256:07e70381bbe801fe9bd8d0add40eacc695cef09910c42eaff4d4190673ef59be

Observation d7737523-5de0-43c3-a758-44fe24b45c1a · outbound

This paper cites mPLUG-DocOwl 1.5: Unified structure learning for OCR- free document understanding.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization mPLUG-DocOwl 1.5: Unified structure learning for OCR- free document understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:14.349022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:03.804645Z digest=sha256:3ade041df2b96df87e06f4e53277d495c5717e51831eefd203ef8acd705a9e76

Observation 9bff7f50-ba8c-4efd-80dd-554c471ec62a · outbound

This paper cites Lora: Low-rank adaptation of large language models.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Lora: Low-rank adaptation of large language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:03.901686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:03.901686Z digest=sha256:372de2f7606ede85db23dae1b2fc1d174e6aa67afe4d3fea48ae0478cea9474a

Observation 6606b3e1-991a-42ff-b0e2-62480fe7f54c · outbound

This paper cites Layoutlmv3: Pre-training for document ai with unified text and image masking.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Layoutlmv3: Pre-training for document ai with unified text and image masking

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:14.125109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:03.994382Z digest=sha256:94d8b139e7bd302aaf2f202fa1d029bde7750c0b76b1e01acb9004344be999b6

Observation 85147d98-064b-4328-8385-1b3d0f26df68 · outbound

This paper cites Icdar2019 compe- tition on scanned receipt ocr and information extraction.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Icdar2019 compe- tition on scanned receipt ocr and information extraction

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:13.931638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:04.108908Z digest=sha256:271c812dd377fd45d9ca573aa0a89f1c634a6b39fe24700a2629add15535e3e0

Observation bc7d70da-177f-46f5-99f3-9c8024905df1 · outbound

This paper cites Funsd: A dataset for form understanding in noisy scanned documents.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Funsd: A dataset for form understanding in noisy scanned documents

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:13.748786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:04.178898Z digest=sha256:4ae7b781b85d3c25fb654b975280e11539eaf6c892238dd83bb08f82d6c85bf1

Observation 320e75f1-9534-4125-b13f-1c8c709c3879 · outbound

This paper cites A diagram is worth a dozen images.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization A diagram is worth a dozen images

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:04.251928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:04.251928Z digest=sha256:ffa6258a7e7c736aba806c58d49b2dc3f42c17fa185fbb1a68673d6e4d67a30e

Observation dfc00b05-2fac-41d0-b892-0692a689a5ee · outbound

This paper cites OCR-free Document Understanding Transformer.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization OCR-free Document Understanding Transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:04.304286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:04.304286Z digest=sha256:0f7468931fbef46f5cb12bca581685fff54f9c66a76ad2091c11b169d3547f13

Observation 6689e449-4978-47f0-ae20-f53b79396df4 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:13.530535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:04.375812Z digest=sha256:142e8740a9d2ab70e0f7bc53267aab1c48035fd9e023929bb02b1b445d0f4da7

Observation e7f6cae0-596e-4cba-adb2-cdd848a8e768 · outbound

This paper cites DocBank: A bench- mark dataset for document layout analysis.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization DocBank: A bench- mark dataset for document layout analysis

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:13.380065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:04.408043Z digest=sha256:8f082caa1f170bbbd40bd5dd7068eb85330c66a3890d90c7f6506d607aaa3625

Observation 9729aeed-e6c8-4829-95ab-5a0d911f373b · outbound

This paper cites Selfdoc: Self-supervised document representation learning.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Selfdoc: Self-supervised document representation learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:13.238668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:04.482077Z digest=sha256:921257a6c280cd00b03120543aa88d99e0ace42b4bd649f796b9842f27468c2f

Observation c7a26764-3cc7-479d-8571-a0fc72c79c92 · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:13.072147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:04.555642Z digest=sha256:b2023fae3175eaddf294775ddda21b167e9e219351970e98d8b788c6ca1efc92

Observation e2683014-4149-4068-9dce-8ebb7e5ef479 · outbound

This paper cites DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:04.645195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:04.645195Z digest=sha256:370ac77f2151ab8e68e0368e9a29e6193f9a7f886aba21f56dd730ca0d56f097

Observation 1ad7ffd0-f94a-426d-80db-5125b053fac6 · outbound

This paper cites Microsoft coco: Common objects in context.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Microsoft coco: Common objects in context

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:04.806022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:04.806022Z digest=sha256:477179adf565c403a22f9cfa2f00404e2640fa8d6f2987d28034fdee6a268a0f

Observation 4d80c291-eed3-44fe-9945-e8fa849f5100 · outbound

This paper cites Feature pyra- mid networks for object detection.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Feature pyra- mid networks for object detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:04.947845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:04.947845Z digest=sha256:104d6a9157b302cc142af7b5b8cd6567fc048dbb6e48b17306fa6cdeb801e54f

Observation e7fb894e-2d7d-48e7-8cc2-7cd6a526d8ac · outbound

This paper cites Visual instruction tuning.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Visual instruction tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:05.026870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:05.026870Z digest=sha256:fc3dd01d5be2406ce64b52e967bb7ff2814eb9fed4bc5fda83cdfd0757c72c38

Observation 02842439-b046-4944-b9c5-5a46b0a4d168 · outbound

This paper cites Improved baselines with visual instruction tuning.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Improved baselines with visual instruction tuning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:12.802843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:05.121906Z digest=sha256:be0302912572d8d027ce633a576c214c02150d0603c683fa6b2f35925d179aa6

Observation cc6f6b2f-255c-40f5-b1df-25125e92583b · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:05.251861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:05.251861Z digest=sha256:32b03396bb90a688a421128f5c3451b9e1b269e37ee64b3e76ba2593b3f00d5f

Observation 691dc397-de21-4031-8702-cf906a213fb4 · outbound

This paper cites Swin transformer v2: Scaling up capacity and resolution.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Swin transformer v2: Scaling up capacity and resolution

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:12.531717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:05.329805Z digest=sha256:69b4e0931dda9f40d822bdd0c97a8af97bff76e55a6564526e947f3ce3f9625a

Observation f052739d-9cc9-4a47-94f9-cabd4510cc10 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:05.416342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:05.416342Z digest=sha256:3bd214878c2a30dd59d269879affdda0413c466521b767e06a56fd82f388c9b0

Observation 7544abc3-553f-4d8b-90f1-89d96377dff3 · outbound

This paper cites Decoupled Weight Decay Regularization.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Decoupled Weight Decay Regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:05.480905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:05.480905Z digest=sha256:a8d6c7842ccff2a40038563c58161f2761c80afb7baae85fb0832e67f1bd2228

Observation 90a8901b-f628-45d3-a1e1-3cf2a4e7b441 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:05.531100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:05.531100Z digest=sha256:45812c03a4b45bb161efaddab95ba2d56e2edb3aff00d08dbdb772147977e633

Observation 40e38e9f-cae0-47d4-9c06-2aae302fb3f7 · outbound

This paper cites DocVQA: A Dataset for VQA on Document Images.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization DocVQA: A Dataset for VQA on Document Images

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:05.622333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:05.622333Z digest=sha256:4b16121f43571088d718c9b6927de864ab8198f6f114e2438f779fb39478b316

Observation 33e1d906-ab40-46a0-bc34-7c2a5cd57de4 · outbound

This paper cites Azure Cognitive Services: Optical Char- acter Recognition (OCR).

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Azure Cognitive Services: Optical Char- acter Recognition (OCR)

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:12.216896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:05.739431Z digest=sha256:61e8067ccd2269be46db07a3191e610fe1c892f5c96c946fe5d00bc55a8ad06d

Observation bc0fdc9f-f52d-4fb0-a1f5-a8e1a0ee34e9 · outbound

This paper cites Rectified linear units im- prove restricted boltzmann machines.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Rectified linear units im- prove restricted boltzmann machines

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:11.930770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:05.831731Z digest=sha256:150b4e9414af3efaf0684104c05a16554b49e3c7e0f9a839f54f090e953e6a5b

Observation dc51e973-6da0-455a-a2b2-dc51d816a9c2 · outbound

This paper cites Cord: a con- solidated receipt dataset for post-ocr parsing.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Cord: a con- solidated receipt dataset for post-ocr parsing

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:11.644166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:05.927075Z digest=sha256:8c52ac6c1f41cf0bbce1f6051ef04b0d6c1a271b6c8085e9ae930d10d5474b0a

Observation 0534d87b-bb01-452c-bb25-29b97a652f07 · outbound

This paper cites Automatic differentiation in pytorch.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Automatic differentiation in pytorch

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:11.417501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:06.001101Z digest=sha256:7aae9ba4c3fba6d462452ed8294497d9ee1196eb9ab3800e8e3882911c04c5b5

Observation 400d844d-0b27-4646-8b4a-c68848708a74 · outbound

This paper cites Doclaynet: A large human- annotated dataset for document-layout segmentation.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Doclaynet: A large human- annotated dataset for document-layout segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:11.171317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:06.059606Z digest=sha256:e25ebda824449e36c831db684ac0543df0ac50f2f7726453b9b6227a97225046

Observation 519d9a20-7fff-4027-8483-af81641f1a3a · outbound

This paper cites Going full-tilt boogie on document understanding with text-image-layout transformer.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Going full-tilt boogie on document understanding with text-image-layout transformer

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:06.153026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:06.153026Z digest=sha256:e99a7fb9e2da69352e58029e799657e3347ca7c27a5d1f3109e4605f17e91986

Observation 5761f58d-16d5-448b-9b94-5b221ebb6dd6 · outbound

This paper cites Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:09.968066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:06.200056Z digest=sha256:61d2e2df00e5cda647636c492ff6ad7406f9f15840ed7140339413915d8f8c80

Observation 5ef67d07-a8dc-41dc-9d26-cb56b53ddcc9 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:06.292975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:06.292975Z digest=sha256:3515cfefc29d85be8f8b959e309967efa5df53fd9c575e38b0e22e6eaffd1c05

Observation f9c86612-5a36-4025-a683-38c90fd824d6 · outbound

This paper cites Deep Learning based Key Information Extraction from Business Documents: Systematic Literature Review.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Deep Learning based Key Information Extraction from Business Documents: Systematic Literature Review

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:59:07.936194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:06.366921Z digest=sha256:10e6c1492aea62e3429f7dc9a348e1398e2a8aa354b17e2ba20c4eae560fb025

Observation 0450fa45-eb9f-42af-997c-a167f0ea5ae5 · outbound

This paper cites Docile benchmark for document information localization and extraction.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Docile benchmark for document information localization and extraction

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:09.303856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:06.430607Z digest=sha256:f43e408ff9303dd279dd61578a91291312e1d965306584eb53e62acc0b407db3

Observation 4f4ff3ee-551a-480a-b2f8-3c623614cef3 · outbound

This paper cites Towards vqa models that can read.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Towards vqa models that can read

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:06.508889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:06.508889Z digest=sha256:f6efb1c20ec0de5baa1c11e7e6df5fa688c8937fdc05278146a1a48d7bd28318

Observation 7927cf82-2d59-4f37-a184-cef2e104c169 · outbound

This paper cites Spatial Dual-Modality Graph Reasoning for Key Information Extraction.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Spatial Dual-Modality Graph Reasoning for Key Information Extraction

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:06.590430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:06.590430Z digest=sha256:54667dc408e9c82a074d10900246df49def6b277fd65883a919dc9d451a91622

Observation abe05721-78e1-4dea-a77f-29a5e9be424e · outbound

This paper cites Instructdoc: A dataset for zero-shot general- ization of visual document understanding with instructions.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Instructdoc: A dataset for zero-shot general- ization of visual document understanding with instructions

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:09.108724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:06.674793Z digest=sha256:ed94dea2419ba4cc95cad7a2b37e10cdae1dc94c96476594c9a6067387f9eec5

Observation 1fda786d-7401-4228-913d-144790c9ec8e · outbound

This paper cites Docllm: A layout-aware gener- ative language model for multimodal document understand- ing.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Docllm: A layout-aware gener- ative language model for multimodal document understand- ing

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:08.900312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:06.763920Z digest=sha256:55ac465ec06211e4ccb9e78fc02ab3012eca79c0f1fa294918fd2b76cc92a1bf

Observation 93db90e9-a191-4f32-86b2-5df0d338957e · outbound

This paper cites Vision-enhanced semantic entity recognition in document images via visually-asymmetric consistency learning.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Vision-enhanced semantic entity recognition in document images via visually-asymmetric consistency learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:08.765536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:06.856658Z digest=sha256:81dbc8fc21a161412fb23de531de8efc6dbddc466e740fea4c19271f84ad6d72

Observation 9212ef6d-d603-452e-9f5f-855e5e9601d3 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:06.956279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:06.956279Z digest=sha256:6afae5faae541de7c31bfd875b891246e7a3e2cc6e4d156bd75a2f86dab68b29

Observation 6292d008-1c23-4547-8c70-31718d9802ac · outbound

This paper cites Layoutlm: Pre-training of text and layout for document image understanding.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Layoutlm: Pre-training of text and layout for document image understanding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:08.562604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:07.028063Z digest=sha256:daff506e4a9ea034b04408156de1d13416b326fe4236d6e1d5223811f2ca0f3c

Observation 86758d95-0187-4e1b-bf58-47ca6d7d7f53 · outbound

This paper cites Layoutlmv2: Multi-modal pre-training for visually-rich document understanding.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Layoutlmv2: Multi-modal pre-training for visually-rich document understanding

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:08.414877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:07.159523Z digest=sha256:9db6d95c6372bdf01aeada8cce804f96dc54d5a2aa9e3f25d22aa7a9e3aa62f3

Observation 2454a794-357c-430c-9d0d-367ccdaf14d4 · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:07.342834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:07.342834Z digest=sha256:58dd93ef3a20db2897abd7649436c3d4068ee28b764c324de2abf7b8509a6be1

Observation 2c7ea86d-7888-4a3b-a0d7-2bf2eeb9d5ad · outbound

This paper cites UReader: Universal OCR-free visually-situated language understand- ing with multimodal large language model.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization UReader: Universal OCR-free visually-situated language understand- ing with multimodal large language model

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:08.280511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:07.428302Z digest=sha256:ecf158f97345a6e313f928fe3fc131a1ec7174853835048226e6683272bc3a64

Observation 249d394c-c77b-4bbf-8550-f94411c7e77b · outbound

This paper cites By my eyes: Grounding multimodal large language models with sensor data via vi- sual prompting.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization By my eyes: Grounding multimodal large language models with sensor data via vi- sual prompting

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:08.118046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:07.466602Z digest=sha256:1604ec8ef5dca5f6b7af87963f5aa8157fd145468be961942153de6860348d62

Observation cd4dc51e-b3e7-4372-80dd-ac741b1daf00 · outbound

This paper cites StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:59:07.773014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:59:07.531504Z digest=sha256:38bfe63c60ae2bdc616a5fed413f321a8e0fc11efabb2b96f7066996305caa98

Observation 1a888602-e291-42c8-9625-1461a92fcc8a · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:07.601054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:07.601054Z digest=sha256:2dd1b6b454184645b6f624b1e51e8050f66654b780f62f44a921be902f640469

Pith citing papers

No inbound Pith citation observations are available.