Pith. sign in

Paper Citation Record · LEDGER

TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2404.09797.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.09797 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T06:06:07.136455Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.324968Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a3550c13-5430-4d16-beb1-4f025fe774c8 · inbound

CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models cites this paper.

CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T06:06:07.136455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T06:06:07.136455Z digest=sha256:24ce25c27fb6e22bc393ca5521fe0bdef010e7a0283735ebcfc7932a63292899

Observation 7428d43e-8f64-49c9-ac41-3279c7a5462a · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.718612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:5bf409155af4a3e9f5236a85ffbc9fa2f356dbd1082a920b0b2d72020f2f62cd

Observation 14d2e6a5-43c5-4b83-9c53-e5550f74b49f · inbound

Training-Free Reasoning and Reflection in MLLMs cites this paper.

Training-Free Reasoning and Reflection in MLLMs TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:55.011386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:55.011386Z digest=sha256:208fa0507648e71270fc0051a48ff619c51be156066dad1ba841e9c1e1ce471e

Observation f2480a51-1e1a-4a8a-a84c-8fb6af04d3b8 · inbound

Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning cites this paper.

Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:40.207197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:40.207197Z digest=sha256:5d689de87b95fccc88ef5a189be2d285723a844bab41f73bd0b2650225ab88b9

Observation 768f4ee9-819d-409a-82a3-6511bcca86bf · inbound

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking cites this paper.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.230798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.230798Z digest=sha256:7bdffbeb416b928bd16f75fff97afe926de4a4278d70c52cd5a130c6e2cf9be6

Observation c86508c6-2314-49f8-aa28-9a37f1f5abea · inbound

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement cites this paper.

Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.313744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.313744Z digest=sha256:a5e5dd299e14b79e369de7f379fd959206b449ed0d41c1ca3014178ad217f37a

Observation c58f9d50-3dee-4194-a524-81aad9ff36cf · inbound

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints cites this paper.

Reinforcing VLMs to Use Tools for Detailed Visual Reasoning Under Resource Constraints TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:50.436455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:57:50.436455Z digest=sha256:1b809b6ddebf0cf916441b5107c5704f7d4f7a06ef8408e9763a385ef7730a29

Observation cf724eff-60da-4cd5-9d17-31b711af7b41 · inbound

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning cites this paper.

VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T16:33:57.441181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:33:57.441181Z digest=sha256:19c96e663c1075c72ffb20129ff3c879a8523386cd54153196872c2fc122c406

Observation 968f33d1-246c-4337-bb3f-7db71a08302a · inbound

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs cites this paper.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:57.265975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:57.265975Z digest=sha256:6adaa84b5a74c71043ca1462b7822a363aafcf80652eb46983439d43aefd0ecd

Observation 956a5d75-d46e-45d8-8fac-b1b44f4ea0a1 · inbound

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding cites this paper.

M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T10:49:58.806693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:49:58.806693Z digest=sha256:d14d18e0cfbcd2509bc98a3ae9c4a267f5411193735c75bf0f63a43266872967

Observation 9a53a9d3-7f3c-4e1b-9c85-eef5ff213a55 · inbound

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework cites this paper.

HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T20:19:15.251753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:19:15.251753Z digest=sha256:1934978188b63e7437e82a7d7928715e4f9b8bc05ddc8cc561ec034e583ce6a5

Observation fc652840-7b6c-4a37-b668-10437765a6d4 · inbound

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward cites this paper.

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:47.884342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T19:59:19.379119Z digest=sha256:a1583963153313e420d7769658889103bb28b4b3933e6d6565e2f0acb223f7da

Observation 9a6ef3ee-0bca-49a3-98e9-2309dbe33b32 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.327372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:5c44e72a797fa6fe7f722e3be43bf19e00964b8a4a682c10e5da6336fefb9f2c