Pith. sign in

Paper Citation Record · LEDGER

TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2403.04473.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.04473 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T23:13:13.913693Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-23T19:43:23.773160Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 541007d1-b508-4c6d-8895-89630c02dfa3 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.414279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:ded2b9f2e078437cbceb3c97a949e03482797588206e03760623eac09406d4ad

Observation dce1f58e-15c5-4d58-9261-e90500800966 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.111646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:1468d1b90fb9ac5585f576d17362e64b464f1e9f60494fcfaac5043925e82b15

Observation 77c8df6e-1d1c-461e-89b0-5599283f556b · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.774762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:ee40757b6b7fef413028a2631e97542e9c026eeaae891496da8147464fd1c1ae

Observation d0d22a99-2dca-4228-b9e2-14d19b081665 · inbound

MiniCPM-V: A GPT-4V Level MLLM on Your Phone cites this paper.

MiniCPM-V: A GPT-4V Level MLLM on Your Phone TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:07:32.096671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T21:07:31.387726Z digest=sha256:8105e9742dfc2f2f7811b386598c5b3a5cee57478242273e54414940fea8da0d

Observation 9f91f9c5-3858-4714-b421-252add1f24f6 · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.827628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:11d41037e1616ce1015097cab4cd45e4a91dd24e381ec36fded46cb7f6c1685d

Observation 3032a2af-cde3-4694-8376-af8e8cfababc · inbound

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model cites this paper.

General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:50:57.878737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:50:57.814634Z digest=sha256:0825dff9af06e53a182c4900a22e81da1a2b46261ae5a1ff4594848ead877a52

Observation 8a11461b-8a19-4f64-b893-9cb44861c518 · inbound

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling cites this paper.

PDF-WuKong: A Large Multimodal Model for Efficient Long PDF Reading with End-to-End Sparse Sampling TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:43:23.776809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T19:39:35.147671Z digest=sha256:c53f2cc71364058f036aea448387b3cba8d2ff0c895626d232d6d1475a0b8f9b

Observation fe01c260-d247-4af5-b785-ac057b13bc24 · inbound

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents cites this paper.

VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T15:37:25.911108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T15:37:25.781240Z digest=sha256:2528a9fc86a28617fd1eaa99b29f6efef3f463a29c6b0df920bfa53453fb4db4

Observation ca16e56a-e147-4132-b09d-d7c651eb04ce · inbound

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction cites this paper.

Document Parsing Unveiled: Techniques, Challenges, and Prospects for Structured Information Extraction TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 142

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:15:47.281080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T19:15:21.695801Z digest=sha256:f85664e13ce22dd43292bc63facd28a703a607bfbd22def4a1ac2c0cc437501c

Observation 7731a63b-db2d-49a7-ae06-be6343d86ea3 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:26.724787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:9ce0386186415a8cfb5fd030e48d0570433e441da62751b597b80ad88d8b66af

Observation b6c8fc8b-a0b1-416b-ad9e-489219011128 · inbound

\'Eclair -- Extracting Content and Layout with Integrated Reading Order for Documents cites this paper.

\'Eclair -- Extracting Content and Layout with Integrated Reading Order for Documents TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T23:13:13.913693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T23:13:13.913693Z digest=sha256:5549415f2f3d8a8e8ed2ab4790105689b0af112667c624bdc6c00eb6359b2991

Observation bf198aa3-25a7-412e-8c04-b444fc36d215 · inbound

EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition cites this paper.

EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T22:57:16.802643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T22:57:16.802643Z digest=sha256:a7678fe71174d1c567d51f8918802b7b3a5a75a45f63f409185d6f982504afbb

Observation 331caa8d-4bf1-46d3-9584-531062fe2fce · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.851273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:baf9e3edc9c9932aee16955800f0855b78d916636cbe58591046931c9efb7e33

Observation 63fe7e2c-b1a3-4b46-9cfb-c91351b1379f · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:26.253023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:26.253023Z digest=sha256:1614124f9ec39a7ee2c01a74fb54c700385b0bb32910847435d2f6a9fb69d304

Observation 0d7adbae-c4c1-44c6-adc2-ce01de7d408e · inbound

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning cites this paper.

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:18.863408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:18.863408Z digest=sha256:aed4cdc2a953e4ca2b805537e5dcf682f973e3b580495992e748e3ca72326436

Observation ab764237-0230-4fb1-8c10-aa1c07c08872 · inbound

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration cites this paper.

Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:52.167671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:52.167671Z digest=sha256:48e8aa4454c1b37a1c5c6bf6c08a54f35c02c96e259cb179061cca26d786a928

Observation afe2df36-233b-4b4b-8ef6-4474a7ff8dfb · inbound

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement cites this paper.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.120767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.120767Z digest=sha256:abf81690de8524faab0b40147ca03009e6d36245ac992501e3ec3736b09e27c4

Observation 55ef876c-a123-4673-beb6-714e832d43db · inbound

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs cites this paper.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:48.010046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:48.010046Z digest=sha256:1d8553baba18d53a5c332c35b0726762d4f42cc61ef60be132b31fd566dfeee7

Observation 0a80537c-96f4-40f2-b0c5-5b88f29a12bf · inbound

Structured Attention Matters to Multimodal LLMs in Document Understanding cites this paper.

Structured Attention Matters to Multimodal LLMs in Document Understanding TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:47:38.535627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:47:38.535627Z digest=sha256:89ea288919d8d7ee242e06e82469e1448ea0bf3dabbca50a6a07a2ae8b79d7e3

Observation cdb269cb-b2e4-4c84-8496-dac5bfbdaa65 · inbound

ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning cites this paper.

ESTR-CoT: Towards Explainable and Accurate Event Stream based Scene Text Recognition with Chain-of-Thought Reasoning TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:39:36.782313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:39:36.782313Z digest=sha256:446272b5eb9afc7c6c54d5a374b11f13c9d5dbe4e1c2daa341f90f6b4f12ff4e

Observation 0a905ad9-f29c-4e9c-b2ec-4f0b1b862264 · inbound

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance cites this paper.

LIRA: Inferring Segmentation in Large Multi-modal Models with Local Interleaved Region Assistance TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:02.762351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:02.762351Z digest=sha256:09cf35cc48149440163dee8072c6726a5cbaa5dd987676dbf30343664fd13ec7

Observation 836f1d85-7cb8-4389-ad5a-7a2d6669fde8 · inbound

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation cites this paper.

Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:43:18.671551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:43:18.671551Z digest=sha256:26103578db01e802c24654320bd1a6aae031ed65aa7f74e327c0e9ba17266d8e

Observation 2195bfaf-1295-49b4-9434-3315e423845d · inbound

Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency cites this paper.

Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:28:11.824348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:28:11.824348Z digest=sha256:00d71c83a937c53f9a24198b97835b44582996d9ae24c6526192f502b870175b

Observation cc6f6b2f-255c-40f5-b1df-25125e92583b · inbound

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization cites this paper.

VDInstruct: Zero-Shot Key Information Extraction via Content-Aware Vision Tokenization TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:05.251861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:59:05.251861Z digest=sha256:a70fe7b81412668208f1f588882335185ff4a5ab559108fa19d1634c855bd251

Observation 146b0fc2-3f5b-499c-9f4e-1565187e2bd5 · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.414900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:f84f496e56538144e3965dd0c7f30692b219ecbf3859831c8cb22c1b2ce33a1b

Observation b2244a65-2661-4e5a-9c3e-68e7cf52ac17 · inbound

Describe Anything Model for Visual Question Answering on Text-rich Images cites this paper.

Describe Anything Model for Visual Question Answering on Text-rich Images TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T16:51:25.051188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:51:25.051188Z digest=sha256:58a94408ce51d127de2134c2e6947aed120b668bce25df61dd800dd1f5c9a320

Observation ece6fc9a-60b6-4784-89f6-ca02fa63b724 · inbound

Docopilot: Improving Multimodal Models for Document-Level Understanding cites this paper.

Docopilot: Improving Multimodal Models for Document-Level Understanding TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:57:02.725020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:57:02.725020Z digest=sha256:b58cf7eb81da888c26aed349aca9b04a4bfb82fdd659dbda6035e547d2cd539e

Observation 1a118deb-9a5b-4c1f-b5ab-fbba38a6e8da · inbound

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification cites this paper.

CogVLA: Cognition-Aligned Vision-Language-Action Model via Instruction-Driven Routing & Sparsification TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T14:42:32.710762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:42:32.710762Z digest=sha256:11d947dc26598ce1ac3222eff53a8947931890a29476b9590a4a339daae4c637

Observation b1e09e38-907f-449a-b072-d1c13ad8f18f · inbound

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing cites this paper.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:25:32.017809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:5dd24da84a982824cfa316f4f15fb35a8215108a79e24068ed5c2dd9372725b4

Observation 97769deb-fd86-43e5-8f63-1cee51cff3d4 · inbound

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters cites this paper.

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T14:16:50.788856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:16:50.788856Z digest=sha256:f2ed7c363f05bf06c71b1d991b1a2cf67b9170af07327778b2d6beea7dcc0e01

Observation a2564aa5-037b-4aff-a5fc-20746eea6fa1 · inbound

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework cites this paper.

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T20:19:14.895125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:19:14.895125Z digest=sha256:ba2fb0604a3e8de1a6ad813fbe1d743adb60f2fc8686386c51a8b6f4d75c66d9

Observation 1e380ce8-469e-41cf-ad6b-d6c707f8fde4 · inbound

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding cites this paper.

Cognitive Mismatch in Multimodal Large Language Models for Discrete Symbol Understanding TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:15:21.113889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T09:11:31.870441Z digest=sha256:4dab4b8011aaa82f96fa055bdf537c09b48d01ea41cf8638975267cfe63ff942

Observation 44327a09-f177-4137-8c73-f66c3ad7f413 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.702294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:5997e4d1c93a088f898a166f4de30fdf11ff799bc3d299f711b6e9f04ab195e2

Observation e2959da3-271d-403f-a559-ddea62488bf4 · inbound

ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment cites this paper.

ReAlign: Optimizing the Visual Document Retriever with Reasoning-Guided Fine-Grained Alignment TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:15:58.250374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:43:15.630570Z digest=sha256:36160e3b7a4402e90f3b812883939cd3c6fe7766c913857c9cbf6f5b3d9b04bd

Observation ddd1ef0b-72b4-4ab3-999f-4f7910dcaa41 · inbound

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding cites this paper.

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:58.472624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:16:58.889065Z digest=sha256:b73a9879509977d68c472a6c621d50a2e0287d313346e7186399b34edcde6348

Observation 16a8f0b1-41c3-4e79-8933-f7731feeafea · inbound

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding cites this paper.

DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:26:24.600822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:17:55.318813Z digest=sha256:141d2e49a3c45ab1b052165be330fde553a48c365242a82b7a485861c1358e5c