Pith. sign in

Paper Citation Record · LEDGER

OCR-free Document Understanding Transformer

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 27 inbound Pith citation observations for arXiv:2111.15664.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.15664 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:06:00.783505Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0287913b-2c9c-4eab-8f96-585f887e7a1f · inbound

Nougat: Neural Optical Understanding for Academic Documents cites this paper.

Nougat: Neural Optical Understanding for Academic Documents OCR-free Document Understanding Transformer

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:42:12.598592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T09:42:12.463309Z digest=sha256:d4e113e24f6ca916d92ed17f50f263a66960c2718d046ebcbf15cdaa4104b92f

Observation 52a57b62-cfb5-49c8-a3d8-3c86fd912529 · inbound

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence cites this paper.

Granite Vision: a lightweight, open-source multimodal model for enterprise Intelligence OCR-free Document Understanding Transformer

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T20:06:00.783505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:06:00.783505Z digest=sha256:e58231f746a33ea599e36ac133a31c0962178023ffd8a7da7851f949f86f1026

Observation ebc87c5a-4786-421f-b717-5b7e78885a89 · inbound

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription cites this paper.

Judge a Book by its Cover: Investigating Multi-Modal LLMs for Multi-Page Handwritten Document Transcription OCR-free Document Understanding Transformer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:25:19.389300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-23T02:23:22.682357Z digest=sha256:72d7b3526e69902001605cd34f9865fcc8e615efe433fb0052808a9509284d59

Observation becc2c97-4015-4431-a148-35d5bfb3a5ab · inbound

Multimodal Tabular Reasoning with Privileged Structured Information cites this paper.

Multimodal Tabular Reasoning with Privileged Structured Information OCR-free Document Understanding Transformer

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:59.939057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:59.939057Z digest=sha256:22838dcbc1e2812947ea5b4e3e4821c7e99bb3fac48015de0081c6d2fa662d3d

Observation f7abf4ff-69a0-4b69-bd96-f571d0398808 · inbound

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs cites this paper.

SparseMM: Head Sparsity Emerges from Visual Concept Responses in MLLMs OCR-free Document Understanding Transformer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:08.514999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:27:08.514999Z digest=sha256:675bacc01bfe741c4c4ce9e1e0d3e12a2347a9b5cef61d1ef569acea095edaca

Observation c3823ade-e530-46d5-a213-ca6e587460e2 · inbound

Digitization of Document and Information Extraction using OCR cites this paper.

Digitization of Document and Information Extraction using OCR OCR-free Document Understanding Transformer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T04:41:40.869980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:41:40.869980Z digest=sha256:7a46a42175073ffd93ae13792bd17a7ce64c19b5649d96898c238d1596413f6f

Observation dfc00b05-2fac-41d0-b892-0692a689a5ee · inbound

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization cites this paper.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization OCR-free Document Understanding Transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:04.304286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:04.304286Z digest=sha256:ef4c6f6566185e6b4a3403748a3683d43b07d7e2075ebfce5a9c56b04ef045b8

Observation 0e9e044b-8a64-4066-8198-02f52f59bb8c · inbound

ExpliCIT-QA: Explainable Code-Based Image Table Question Answering cites this paper.

ExpliCIT-QA: Explainable Code-Based Image Table Question Answering OCR-free Document Understanding Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:15.377045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:15.377045Z digest=sha256:aaacf253e9d6d1d38b2985130ab421d1b948abb77e8406921d98608d07e9d5af

Observation f1a1c0d5-70ad-4be3-b68f-9c95e90c72a5 · inbound

FRED: Financial Retrieval-Enhanced Detection and Editing of Hallucinations in Language Models cites this paper.

FRED: Financial Retrieval-Enhanced Detection and Editing of Hallucinations in Language Models OCR-free Document Understanding Transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:36.844593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T13:10:36.844593Z digest=sha256:1fda3a990d30843c7249641d46e376ecfdb79018d7ce26835af79ba22dfcb90d

Observation 4cb40f2c-26a4-4ef4-9e94-86a243878492 · inbound

CF-RAG: A Dataset and Method for Carbon Footprint QA Using Retrieval-Augmented Generation cites this paper.

CF-RAG: A Dataset and Method for Carbon Footprint QA Using Retrieval-Augmented Generation OCR-free Document Understanding Transformer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:49.553248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:49.553248Z digest=sha256:9b3d69633231d29dd06e753317ba777c0c1b3d2c9ca07528fe646c748f0a77cb

Observation fd652515-a3b9-4feb-980c-ac52d70ae8b2 · inbound

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition cites this paper.

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition OCR-free Document Understanding Transformer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:52:12.735088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:52:12.735088Z digest=sha256:b88e1ea25740251538dc29dd4ea73cec319c795e212bc9908e8deec7fd1123e0

Observation b90a1607-35fa-4663-8aea-a5227db421bc · inbound

Reverse Browser: Vector-Image-to-Code Generator cites this paper.

Reverse Browser: Vector-Image-to-Code Generator OCR-free Document Understanding Transformer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T05:46:40.982698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:46:40.982698Z digest=sha256:cfbf27896569ba7527668730075b082f7a9f28d0a651029026fe2f8b275b51b4

Observation 5b20f480-1a5c-4e64-9d70-00bddd68daa2 · inbound

CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding cites this paper.

CodeOCR: On the Effectiveness of Vision Language Models in Code Understanding OCR-free Document Understanding Transformer

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:32:36.459987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T08:30:50.984873Z digest=sha256:3d29bc4d53271fc61cc14ba275ecc2d488c92d8b02586804fc84e42b7675d100

Observation 3e885d8f-f70b-4630-badb-5dedd7b558ac · inbound

From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models cites this paper.

From Plausibility to Verifiability: Risk-Controlled Generative OCR with Vision-Language Models OCR-free Document Understanding Transformer

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:55:19.416250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T08:53:18.268970Z digest=sha256:b2a28dd1ad3c0712513f4a47c3b85ae19c0d7441f8d684bfc0f60ba81b57fc69

Observation 9b5a0557-389f-4878-8a9a-7bc6cdcc713e · inbound

Fine-tuning DeepSeek-OCR-2 for Molecular Structure Recognition cites this paper.

Fine-tuning DeepSeek-OCR-2 for Molecular Structure Recognition OCR-free Document Understanding Transformer

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T19:38:10.451189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T19:36:45.799571Z digest=sha256:8f9b62fb578ce6473fcb57709d54035e21613d510179290ba73f1cb27bb6e3cc

Observation 77aeb6da-aaeb-4584-b266-b456e2b3d427 · inbound

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale cites this paper.

MinerU2.5-Pro: Pushing the Limits of Data-Centric Document Parsing at Scale OCR-free Document Understanding Transformer

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:35:52.523389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:58:41.377996Z digest=sha256:c8f197da323adbfbdf308995ffe37ea5891af8495ffa6bad871264c5cb81d5da

Observation b952bad2-f67a-4d2d-b2ea-a024469d834a · inbound

From Handwriting to Structured Data: Benchmarking AI Digitisation of Handwritten Forms cites this paper.

From Handwriting to Structured Data: Benchmarking AI Digitisation of Handwritten Forms OCR-free Document Understanding Transformer

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:02.232451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:08:30.056580Z digest=sha256:85b318675450deb4d99224df983e8005afe256bf5c2de02134e12472c392879a

Observation 9dc05924-4aa7-4363-944a-68b82b6681e8 · inbound

MADP: A Multi-Agent Pipeline for Sustainable Document Processing with Human-in-the-Loop cites this paper.

MADP: A Multi-Agent Pipeline for Sustainable Document Processing with Human-in-the-Loop OCR-free Document Understanding Transformer

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:28:21.549307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T14:25:11.234515Z digest=sha256:a5093cffd8949437351ea2355a84df7e787f7a9bfcbbf0ebb7829663dbdccfe8

Observation 3c742d3f-4367-4d74-8339-82b4d9043148 · inbound

FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing cites this paper.

FastOCR: Dynamic Visual Fixation via KV Cache Pruning for Efficient Document Parsing OCR-free Document Understanding Transformer

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:58:24.803779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T14:53:57.376715Z digest=sha256:2fbf5012b6e839799fc18962214cfade1100d3782d4194650d6c83b89a4ddce7

Observation 079d3557-8c18-4003-a389-82f187a7b1db · inbound

Structured Layout Priors for Robust Out-of-Distribution Visual Document Understanding cites this paper.

Structured Layout Priors for Robust Out-of-Distribution Visual Document Understanding OCR-free Document Understanding Transformer

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:05.051758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:48:34.771799Z digest=sha256:78d536e68bf4645a55b383aa870a9da6ac0e82b31f8f0678cb48f1b5fb907f52

Observation 8603ba34-9d37-4353-ab84-7fd608dc9c14 · inbound

Multimodal Approaches for Visually-Rich Document Type Classification: A Comparative Analysis cites this paper.

Multimodal Approaches for Visually-Rich Document Type Classification: A Comparative Analysis OCR-free Document Understanding Transformer

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:46:19.656415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T15:03:16.112327Z digest=sha256:ca9df109939333b8b11c930c4a75c1f0b99b7f16dcbeb39e431af86f4ef0b54f

Observation d7a29775-8901-43c5-8695-8194ce2dd7aa · inbound

TrOCR for Medieval HTR: A Systematic Ablation Study with Cross-Dataset Validation cites this paper.

TrOCR for Medieval HTR: A Systematic Ablation Study with Cross-Dataset Validation OCR-free Document Understanding Transformer

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:09:57.074839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T00:54:53.740294Z digest=sha256:3fa2a4f07e2cfd5d8731e6cc5ffadab195b1be47113b8ca2b3fe053de349a453

Observation a19f4980-2e49-4b25-9909-6bb977e9cc9e · inbound

LLM-Based Examination of Eligibility Criteria from Securities Prospectuses at the German Central Bank cites this paper.

LLM-Based Examination of Eligibility Criteria from Securities Prospectuses at the German Central Bank OCR-free Document Understanding Transformer

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T03:58:57.095353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T03:56:28.271760Z digest=sha256:2afa19cd54656eaba7c06914925c279d6d27d1f0ec42b0eb219aeb0e337599b0

Observation b55c7121-1f6a-4218-8031-501c006fc1bb · inbound

Causal Connections: Leveraging Multilingual Fine-Tuning for Financial QA@FinCausal 2026 cites this paper.

Causal Connections: Leveraging Multilingual Fine-Tuning for Financial QA@FinCausal 2026 OCR-free Document Understanding Transformer

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T02:23:00.982572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T02:17:26.872854Z digest=sha256:f41a7aad279f3a87707f50424a775556ab7d4c9c969297d35e463f77057a86b5

Observation 751d240c-522c-4644-90d7-b6a264841737 · inbound

Scalable Visual Pretraining for Language Intelligence cites this paper.

Scalable Visual Pretraining for Language Intelligence OCR-free Document Understanding Transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T07:33:07.528113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:33:07.528113Z digest=sha256:c7db11f62c60afb253a617bf3d84c283b91d819b5cc9be40965aa00b1b05d3c6

Observation 8533fd70-5f9f-438a-9e3e-d0b5c98c0480 · inbound

Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images cites this paper.

Pixels for Programs? A Cross-Provider Case Study of Input-Token Accounting for Source Code as Text and Images OCR-free Document Understanding Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T08:35:42.265776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:35:42.265776Z digest=sha256:5b9988e4b20099580728c4342e7000a134b01a64684589ce63eb34cfd1a50f0e

Observation d9228e72-7654-4318-ae76-812b1bc256f4 · inbound

XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding cites this paper.

XL-DocBench: Benchmarking Evidence-Grounded Extra-Long Document Understanding OCR-free Document Understanding Transformer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T01:39:07.682081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:39:07.682081Z digest=sha256:cd7cbe1590cc1c4efedc214ccf16591559fe02ffd89a1f9bab63c5a57e2ebd7a