Pith. sign in

Paper Citation Record · LEDGER

TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2410.05261.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.05261 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:31.079587Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T07:13:06.644583Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a6719a44-8107-49c8-ba88-5c2df4e4c016 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 286

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.272883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:6eef51fa588ecd4d3afbf297ff1e925da8a8f36267c8a275905559871e99232f

Observation 51d4c1cb-6033-4345-a32b-c74ed19ae48d · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:09:26.325640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:5e426b7a95efd073e1b5b4d9cc62c58380ae00daa37994781037ae4566ee00ac

Observation 8e87d110-f4fe-44ac-9097-1b8a3435c0db · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 141

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.163440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:71bb3961e23a9f55a7d468a8fa2eb43fe48154181305b1bb337422b136e910c0

Observation d579bcc5-90ce-4206-a9d1-ca44152c45a2 · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:31.079587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:31.079587Z digest=sha256:5ac7e5197857d11844d409b0fbf08d043d705c0f49a789f5695757edc1a8593a

Observation 73163043-530e-41f2-8ede-891deaf0a7d3 · inbound

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning cites this paper.

OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:21.429164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:21.429164Z digest=sha256:2429307bda4792f3f788b4add24c370f6c0adab76384e42f8d852743a6c08b4a

Observation 1b75d618-2b7d-4489-bec3-61f553f0e1ad · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.361196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:ace09e042da676bd2ab639409e671a99cebe4906529a46b068c9ac7502aadb58

Observation 2ddefe12-f8f7-430b-932b-7aec5aeb1566 · inbound

Multi-Agent Interactive Question Generation Framework for Long Document Understanding cites this paper.

Multi-Agent Interactive Question Generation Framework for Long Document Understanding TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:45:18.536837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:45:18.536837Z digest=sha256:470aa2ff90021713a8feb17e3d8c9de6645a93aca37a2c45accde0e748ed5e7e

Observation 04693b41-5900-4236-b76b-41a5c210de4b · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:58.933332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:839ffe0b7cea121d58ef72fdf3b851d9f70c088c86267e321a9172beb21d39ea

Observation 81fefa0e-6857-40a4-adf0-e7f5603a0335 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.644288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:886f73f0a92c58f7f57b1dc902d22fc11098807ca4ed9901af8e75e675f2ba23

Observation 19b594b1-1a15-4592-a419-388cb0803ad0 · inbound

UIPress: Bringing Optical Token Compression to UI-to-Code Generation cites this paper.

UIPress: Bringing Optical Token Compression to UI-to-Code Generation TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:00:59.204276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:21:32.024105Z digest=sha256:df29a6ea76b57bc635669ce8894bd40579dd825cf15bb8024023769f458b968d

Observation 793c502c-88b8-4fad-ab15-0882408abbc4 · inbound

Visual Preference Optimization with Rubric Rewards cites this paper.

Visual Preference Optimization with Rubric Rewards TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:00.866790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:45:52.980881Z digest=sha256:c82e65c45130578b62915028c471822a93d1b3d4fdaccfe087131ca60830775d

Observation daf5c038-6a87-4bd2-ad6e-2659742359a8 · inbound

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems cites this paper.

MetaRA: Metamorphic Robustness Assessment for Multimodal Large Language Model-based Visual Question Answering Systems TextHawk2: A Large Vision-Language Model Excels in Bilingual OCR and Grounding with 16x Fewer Tokens

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:13:06.646603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:08:27.096029Z digest=sha256:ac197faaec91c9be404bbdef0a44f28e03a117ee885d886cc6a6df9095f80b49