Pith. sign in

Paper Citation Record · LEDGER

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization

As of 8 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2507.09531.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09531 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:59:07.601054Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 29dfd4c0-64c6-4a2c-acdd-19fef097dacb · outbound

This paper cites Docformer: End-to-end transformer for document understanding.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Docformer: End-to-end transformer for document understanding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:15.872435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:02.282868Z digest=sha256:1f35043a799b6ecfe8c566c3eb34d6501489a6531254caca9146743754966e81

Observation 5508c845-f007-4386-b509-bdb252ee4268 · outbound

This paper cites Qwen-vl: A versatile vision-language model for un- derstanding, localization.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Qwen-vl: A versatile vision-language model for un- derstanding, localization

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:15.717043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:02.385870Z digest=sha256:d93edcebddaf1807a49688e73a6283d228db4095e371b276176733be08833f09

Observation 39ece4dc-3400-49d9-a18a-1fba010b1dc6 · outbound

This paper cites Due: End-to-end document understand- ing benchmark.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Due: End-to-end document understand- ing benchmark

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:15.610275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:02.495887Z digest=sha256:cd86530fc0264b0b4d084af65b97345de624ebe26cb51b4b3ede253de8101bf3

Observation e0ffee70-b30d-4c9a-a799-f5fae76f8f5d · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:02.588039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:02.588039Z digest=sha256:70ad9428b5a374bc7eac7363ae41f70b271e85c836882ff2e8087b19767b74c6

Observation be0e5316-1b31-4aa0-94eb-a6704667ce56 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Gonzalez, Ion Stoica, and Eric P

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:15.523868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:02.710552Z digest=sha256:242695a411e4c74d587a713f5912102cef247374daba9a1a8890fb91ab61cf1c

Observation ad2367e2-b82d-45f9-8574-151d13ee6d98 · outbound

This paper cites Instructblip: towards general- purpose vision-language models with instruction tuning.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Instructblip: towards general- purpose vision-language models with instruction tuning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:15.377660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:02.852460Z digest=sha256:d651dec21f007c422025db185a3a124640954e12a50c0065d98edb710ef5f599

Observation 2562b1e0-86bb-4dd2-b276-104e062ac7ab · outbound

This paper cites Docpedia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Docpedia: Unleashing the power of large multimodal model in the frequency domain for versatile document understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:15.224822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:03.002811Z digest=sha256:fc8bef5b1aa331c9d59a7d228bce43e9dbce1f53160f834e9b26fed50eb2a8f4

Observation 72ea372f-a773-49ab-b868-f35ea1b0ddd1 · outbound

This paper cites Unidoc: Unified pretraining framework for document understanding.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Unidoc: Unified pretraining framework for document understanding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:15.072082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:03.160371Z digest=sha256:2586f81728f204be54b6daea143f827fdf9b51ea6df477d0e9b9fd6c07a4847c

Observation e1a838da-28c7-4965-8f29-5a4561e6547e · outbound

This paper cites Deep residual learning for image recognition.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Deep residual learning for image recognition

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:14.952677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:03.291969Z digest=sha256:83321a1bbb01dd027ae22c62510568f3e39e401e999dd0d3c5bf4823c0a4c565

Observation b584473d-7079-4be1-ace2-70cadbd24256 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Gaussian Error Linear Units (GELUs)

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:03.420452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:03.420452Z digest=sha256:849169f56d423443beeb2afa385f02348a7ed43fb4e45d17ee573d939490e233

Observation 1fec0b0f-ce6d-4e54-ad0f-301f5e9fcad5 · outbound

This paper cites Cogagent: A visual language model for gui agents.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Cogagent: A visual language model for gui agents

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:14.780291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:03.568961Z digest=sha256:3448edc49447d0134f36681544fd09969d4497557fe8378bd553cc1aa403877e

Observation 33ee729d-cb83-4a84-8245-2eb18fe80198 · outbound

This paper cites SciCap: Generating captions for scientific figures.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization SciCap: Generating captions for scientific figures

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:14.636710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:03.706824Z digest=sha256:0c1a673316ff79883339b994c5a4c0e9604987b9c2f3653af4a9e964915894d3

Observation d7737523-5de0-43c3-a758-44fe24b45c1a · outbound

This paper cites mPLUG-DocOwl 1.5: Unified structure learning for OCR- free document understanding.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization mPLUG-DocOwl 1.5: Unified structure learning for OCR- free document understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:14.349022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:03.804645Z digest=sha256:c868a8c7fcd9cf5153a85b8dd2d7943b33692921226b8330cb3ff0da738fe842

Observation 9bff7f50-ba8c-4efd-80dd-554c471ec62a · outbound

This paper cites Lora: Low-rank adaptation of large language models.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Lora: Low-rank adaptation of large language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:03.901686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:03.901686Z digest=sha256:992a140e5fc5832d0b67aabe0abbd8e4b71ec32c0fd6a735ff3a75f48704f761

Observation 6606b3e1-991a-42ff-b0e2-62480fe7f54c · outbound

This paper cites Layoutlmv3: Pre-training for document ai with unified text and image masking.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Layoutlmv3: Pre-training for document ai with unified text and image masking

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:14.125109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:03.994382Z digest=sha256:5109dafd442179269bce41b5445761fea1e01e8f41ff4b35d84b31adbe148481

Observation 85147d98-064b-4328-8385-1b3d0f26df68 · outbound

This paper cites Icdar2019 compe- tition on scanned receipt ocr and information extraction.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Icdar2019 compe- tition on scanned receipt ocr and information extraction

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:13.931638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:04.108908Z digest=sha256:22cdbc0ef21ee7f867e927e704f49a92597cae6c17f4f4bf0c3e28c8690ac831

Observation bc7d70da-177f-46f5-99f3-9c8024905df1 · outbound

This paper cites Funsd: A dataset for form understanding in noisy scanned documents.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Funsd: A dataset for form understanding in noisy scanned documents

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:13.748786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:04.178898Z digest=sha256:b28bd4ba3eea2674e07ac0e727593219e45f7f0c189ff5b09335f16ff6f0fe55

Observation 320e75f1-9534-4125-b13f-1c8c709c3879 · outbound

This paper cites A diagram is worth a dozen images.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization A diagram is worth a dozen images

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:04.251928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:04.251928Z digest=sha256:642fe2ecae455432f965f1ee578e13f06a5cdc6a46be3d5b31c1dae90af0d01a

Observation dfc00b05-2fac-41d0-b892-0692a689a5ee · outbound

This paper cites OCR-free Document Understanding Transformer.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization OCR-free Document Understanding Transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:04.304286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:04.304286Z digest=sha256:15006c045a9f9def90a6ece6d83d264b43f1396f36bcdde497372abdc35282cf

Observation 6689e449-4978-47f0-ae20-f53b79396df4 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:13.530535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:04.375812Z digest=sha256:a182d047e8a74d6e6583a3a27d11cd9eaff52edccbafa6bdf490a2613a7e9f0e

Observation e7f6cae0-596e-4cba-adb2-cdd848a8e768 · outbound

This paper cites DocBank: A bench- mark dataset for document layout analysis.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization DocBank: A bench- mark dataset for document layout analysis

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:13.380065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:04.408043Z digest=sha256:9f6560943e3ca3850ce9dda484b87e4e72fe5a3cc9b3caa92e585ad9f5feade2

Observation 9729aeed-e6c8-4829-95ab-5a0d911f373b · outbound

This paper cites Selfdoc: Self-supervised document representation learning.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Selfdoc: Self-supervised document representation learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:13.238668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:04.482077Z digest=sha256:88c585b74d7058908b791fbc9726471b12ec72eacb7ba756d44a9451743c2275

Observation c7a26764-3cc7-479d-8571-a0fc72c79c92 · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:13.072147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:04.555642Z digest=sha256:7b6678759813a999f0668ad0fddf3179b75192651bc2a0cea1da6699e86d9f34

Observation e2683014-4149-4068-9dce-8ebb7e5ef479 · outbound

This paper cites DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization DocLayLLM: An Efficient Multi-modal Extension of Large Language Models for Text-rich Document Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:04.645195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:04.645195Z digest=sha256:27b29824ff4ff386a57552e11e060598ecfc240ded8680932c5b6f8c819ddbdd

Observation 1ad7ffd0-f94a-426d-80db-5125b053fac6 · outbound

This paper cites Microsoft coco: Common objects in context.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Microsoft coco: Common objects in context

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:04.806022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:04.806022Z digest=sha256:84f82e48ff345492951eae5301df1126820ca8d2bcc11e38438c8f33a378cffa

Observation 4d80c291-eed3-44fe-9945-e8fa849f5100 · outbound

This paper cites Feature pyra- mid networks for object detection.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Feature pyra- mid networks for object detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:04.947845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:04.947845Z digest=sha256:c76ae1066b85989617b13ca0ddfeb205063122ef2f4103ab756f7814251ff260

Observation e7fb894e-2d7d-48e7-8cc2-7cd6a526d8ac · outbound

This paper cites Visual instruction tuning.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Visual instruction tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:05.026870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:05.026870Z digest=sha256:03279dd0841d0a39add57547bd9bbdc142b7b2b27aab55df6bcc6623cf46e4bc

Observation 02842439-b046-4944-b9c5-5a46b0a4d168 · outbound

This paper cites Improved baselines with visual instruction tuning.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Improved baselines with visual instruction tuning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:12.802843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:05.121906Z digest=sha256:689f810afddac12ecb2081cd2d11ad51ae632e8e2ded8adc008006810df522a3

Observation cc6f6b2f-255c-40f5-b1df-25125e92583b · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:05.251861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:05.251861Z digest=sha256:e3c6ede40bcf98893aa2d27709eafe44da39645223165c5ec5c31ace643935cc

Observation 691dc397-de21-4031-8702-cf906a213fb4 · outbound

This paper cites Swin transformer v2: Scaling up capacity and resolution.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Swin transformer v2: Scaling up capacity and resolution

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:12.531717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:05.329805Z digest=sha256:aee5e35cd33e3b39b96b5e0c2ffe3bfe90a23da64c012895999323d544baabd7

Observation f052739d-9cc9-4a47-94f9-cabd4510cc10 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:05.416342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:05.416342Z digest=sha256:f9a889ac2e41f64521dca3f202ae3cc853c54c0e545d1a35045fa174a4a02e4d

Observation 7544abc3-553f-4d8b-90f1-89d96377dff3 · outbound

This paper cites Decoupled Weight Decay Regularization.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Decoupled Weight Decay Regularization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:05.480905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:05.480905Z digest=sha256:0fed384aa1b8700f6ece3e53b85b8b7d2f8eb9e6a81ef9d020fc247acec08a57

Observation 90a8901b-f628-45d3-a1e1-3cf2a4e7b441 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:05.531100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:05.531100Z digest=sha256:8b853e0e87d6a949918cf469b6022a43f0721a6dc664e83231376b9e5fd784e0

Observation 40e38e9f-cae0-47d4-9c06-2aae302fb3f7 · outbound

This paper cites DocVQA: A Dataset for VQA on Document Images.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization DocVQA: A Dataset for VQA on Document Images

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:05.622333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:05.622333Z digest=sha256:5f2de58572f98203205257f2544d4bb7d419c7e3780fa14ef2fb2f4a66f528de

Observation 33e1d906-ab40-46a0-bc34-7c2a5cd57de4 · outbound

This paper cites Azure Cognitive Services: Optical Char- acter Recognition (OCR).

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Azure Cognitive Services: Optical Char- acter Recognition (OCR)

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:12.216896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:05.739431Z digest=sha256:6ce438fa5d53f8cde6e7aa24353911a2e2412ace8fd06540bb6272311539f489

Observation bc0fdc9f-f52d-4fb0-a1f5-a8e1a0ee34e9 · outbound

This paper cites Rectified linear units im- prove restricted boltzmann machines.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Rectified linear units im- prove restricted boltzmann machines

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:11.930770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:05.831731Z digest=sha256:9e5f3593198c388b61a14936175302195696507b4c8c08b505bb8f334d369f8a

Observation dc51e973-6da0-455a-a2b2-dc51d816a9c2 · outbound

This paper cites Cord: a con- solidated receipt dataset for post-ocr parsing.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Cord: a con- solidated receipt dataset for post-ocr parsing

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:11.644166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:05.927075Z digest=sha256:60690c96f04efb895331692d64abec45cc3e8a2c5447fc2d08a4961f7371dc8e

Observation 0534d87b-bb01-452c-bb25-29b97a652f07 · outbound

This paper cites Automatic differentiation in pytorch.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Automatic differentiation in pytorch

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:11.417501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:06.001101Z digest=sha256:d1906fc5f6fce2cffd52d6642d6098357456c82e8a89284dae2b401b68488761

Observation 400d844d-0b27-4646-8b4a-c68848708a74 · outbound

This paper cites Doclaynet: A large human- annotated dataset for document-layout segmentation.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Doclaynet: A large human- annotated dataset for document-layout segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:11.171317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:06.059606Z digest=sha256:91df5df561968bd12b733b419f4f6893634aab71d24e3ba730722dd6e6ca032a

Observation 519d9a20-7fff-4027-8483-af81641f1a3a · outbound

This paper cites Going full-tilt boogie on document understanding with text-image-layout transformer.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Going full-tilt boogie on document understanding with text-image-layout transformer

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:06.153026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:06.153026Z digest=sha256:3ac88ec4665142497c11dbc2f67b01609a37191cbdd38d33144fe03d3fc64a39

Observation 5761f58d-16d5-448b-9b94-5b221ebb6dd6 · outbound

This paper cites Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Deepspeed: System optimizations enable train- ing deep learning models with over 100 billion parameters

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:09.968066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:06.200056Z digest=sha256:c707ee8034b02c074f2b94a53268842ae088370986c66ad1ea0f886fd05650cf

Observation 5ef67d07-a8dc-41dc-9d26-cb56b53ddcc9 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:06.292975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:06.292975Z digest=sha256:397bfccf05c5baceda6026c93081491164e37ae03c0d6dea659615ea9db682e9

Observation f9c86612-5a36-4025-a683-38c90fd824d6 · outbound

This paper cites Deep Learning based Key Information Extraction from Business Documents: Systematic Literature Review.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Deep Learning based Key Information Extraction from Business Documents: Systematic Literature Review

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:59:07.936194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:06.366921Z digest=sha256:61880151a953e5b1acec048023f314dd70907d00b285ec42a4dc8258a0abb020

Observation 0450fa45-eb9f-42af-997c-a167f0ea5ae5 · outbound

This paper cites Docile benchmark for document information localization and extraction.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Docile benchmark for document information localization and extraction

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:09.303856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:06.430607Z digest=sha256:48955777e81d1f28bfa2c970bfb957e66ff79f41c0333ede754ce426997a1744

Observation 4f4ff3ee-551a-480a-b2f8-3c623614cef3 · outbound

This paper cites Towards vqa models that can read.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Towards vqa models that can read

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:06.508889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:06.508889Z digest=sha256:563689da29e23c0f121e37fc182b850dc33a7009d6e0bf3d89104f58dd233020

Observation 7927cf82-2d59-4f37-a184-cef2e104c169 · outbound

This paper cites Spatial Dual-Modality Graph Reasoning for Key Information Extraction.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Spatial Dual-Modality Graph Reasoning for Key Information Extraction

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:06.590430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:06.590430Z digest=sha256:d54d13a7e3d386eb440cb960bf1d25e654f258c8afe717adfcc14f9c95a674ce

Observation abe05721-78e1-4dea-a77f-29a5e9be424e · outbound

This paper cites Instructdoc: A dataset for zero-shot general- ization of visual document understanding with instructions.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Instructdoc: A dataset for zero-shot general- ization of visual document understanding with instructions

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:09.108724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:06.674793Z digest=sha256:bca455640257a6053beb15dc1d35ca1e0f9c6cd2969804e3f0d0d385a12f57bc

Observation 1fda786d-7401-4228-913d-144790c9ec8e · outbound

This paper cites Docllm: A layout-aware gener- ative language model for multimodal document understand- ing.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Docllm: A layout-aware gener- ative language model for multimodal document understand- ing

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:08.900312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:06.763920Z digest=sha256:d3d3a2510627369adc86428ffd93f89d81d07f8fc1ab78b38cc5f371b82b4c4b

Observation 93db90e9-a191-4f32-86b2-5df0d338957e · outbound

This paper cites Vision-enhanced semantic entity recognition in document images via visually-asymmetric consistency learning.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Vision-enhanced semantic entity recognition in document images via visually-asymmetric consistency learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:08.765536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:06.856658Z digest=sha256:c7994c37ff0ff98fbef023aec11fecffba168ef94eae963f2aff743f4462d575

Observation 9212ef6d-d603-452e-9f5f-855e5e9601d3 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:06.956279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:06.956279Z digest=sha256:456ae6d3278274016f28ed5f84a6b34e279ad6ed7fffac8b7b5ddf60e1bbaa00

Observation 6292d008-1c23-4547-8c70-31718d9802ac · outbound

This paper cites Layoutlm: Pre-training of text and layout for document image understanding.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Layoutlm: Pre-training of text and layout for document image understanding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:08.562604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:07.028063Z digest=sha256:723ea947ef91baebb2895fca377cb6c5e5500a3bac81ed10998c732826c9b7fc

Observation 86758d95-0187-4e1b-bf58-47ca6d7d7f53 · outbound

This paper cites Layoutlmv2: Multi-modal pre-training for visually-rich document understanding.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization Layoutlmv2: Multi-modal pre-training for visually-rich document understanding

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:08.414877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:07.159523Z digest=sha256:e3e8067966f9c8adbc401ebdb41f508dd4515be2c3cbdfad5c94e44bfdc10a53

Observation 2454a794-357c-430c-9d0d-367ccdaf14d4 · outbound

This paper cites mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:07.342834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:07.342834Z digest=sha256:3161624ad44868832be7704edaf23b556b42fe210325eed95f64c035516deff5

Observation 2c7ea86d-7888-4a3b-a0d7-2bf2eeb9d5ad · outbound

This paper cites UReader: Universal OCR-free visually-situated language understand- ing with multimodal large language model.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization UReader: Universal OCR-free visually-situated language understand- ing with multimodal large language model

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:08.280511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:07.428302Z digest=sha256:21eba0f666b12065166ff2e372e060a440aaa63a5fb7b98fbe7a58c29c288900

Observation 249d394c-c77b-4bbf-8550-f94411c7e77b · outbound

This paper cites By my eyes: Grounding multimodal large language models with sensor data via vi- sual prompting.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization By my eyes: Grounding multimodal large language models with sensor data via vi- sual prompting

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:59:08.118046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:07.466602Z digest=sha256:06290e0ea14683d224377e0b2359c7d35224e8e6af5c7c1751decb0a43ca3d8c

Observation cd4dc51e-b3e7-4372-80dd-ac741b1daf00 · outbound

This paper cites StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization StrucTexTv2: Masked Visual-Textual Prediction for Document Image Pre-training

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:59:07.773014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T17:59:07.531504Z digest=sha256:8b88b380b24baea77a14aa42204dc012be2413bf5b3b1c111cfc6a17406c2f47

Observation 1a888602-e291-42c8-9625-1461a92fcc8a · outbound

This paper cites LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization LLaVAR: Enhanced Visual Instruction Tuning for Text-Rich Image Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:07.601054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:07.601054Z digest=sha256:98693cbc430ba81216cc6887bc69691feea2062f9b8a0e25e8a03ee2c2f38f2a

Pith citing papers

No inbound Pith citation observations are available.