Pith. sign in

Paper Citation Record · LEDGER

A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2407.01976.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.01976 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:14:20.420778Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T15:28:33.657485Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 09d899ac-7f64-4ac2-a3d6-dab735b2d219 · inbound

Cross-Modal Synergies: Unveiling the Potential of Motion-Aware Fusion Networks in Handling Dynamic and Static ReID Scenarios cites this paper.

Cross-Modal Synergies: Unveiling the Potential of Motion-Aware Fusion Networks in Handling Dynamic and Static ReID Scenarios A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T18:14:20.420778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:14:20.420778Z digest=sha256:c5afb430b9b57671a3bfacb7356687ed02e21ca14e4d76ab35de64d6e3220d90

Observation 03bcc8a6-0566-44e1-a561-5d4d8e92ea4d · inbound

Trajectory Map-Matching in Urban Road Networks Based on RSS Measurements cites this paper.

Trajectory Map-Matching in Urban Road Networks Based on RSS Measurements A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-09T15:56:19.619968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:56:19.619968Z digest=sha256:a94c7bf50badbc8c838e524ee9ddcba741c0f9906f0e0b987f42b306b361e614

Observation cb69cae7-253d-40a6-bd3a-2270f9233f93 · inbound

Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning cites this paper.

Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:12.320432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:12.320432Z digest=sha256:e886a7d37784455bbea4d091fbb1cd457693771e9198a9c69c31a5caf4c25f3b

Observation 85a5e8bd-b95a-43f0-bd80-3bf9d9791459 · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.419842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:6725b16fbb658a9fbd5a1c04dd3c2d12fcaf298a368a6607add1290fca319dc5

Observation cba0332e-a9af-4a0d-a7fd-464e56c6f492 · inbound

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA cites this paper.

DocVAL: Validated Chain-of-Thought Distillation for Grounded Document VQA A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:44:02.464751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T04:41:55.030673Z digest=sha256:f4b26980928fa6f85291dea2fabeb98709f38bccd7cdeeb4dfd30617764d7865

Observation e88c8bbc-22ac-4f14-82e9-0c283b75e26c · inbound

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data cites this paper.

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:50.163129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:50.163129Z digest=sha256:15affd8b92422c011aa4d16be2e7bd08b7119facfa32bf2d9bdfc79ac9f3bf17

Observation b3c6f218-19a5-4075-bf0f-11823da03cf9 · inbound

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method cites this paper.

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:42.102598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:42.102598Z digest=sha256:6f4a95b88c2e786026209f13de84541faf6af63df7fb09a81a3543e7fa9e8c79

Observation 713fc472-408b-4ce2-809d-36530bae406e · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.628581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:14562e393c31a7a85fbd8f15a25c5f0911299c4702ad414e88abee8a25c27153

Observation 3d9e7ac5-b4bc-4469-b036-8ab06e136744 · inbound

DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning cites this paper.

DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:04:21.422148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:01:20.803078Z digest=sha256:1b16dc47fac9c62cc3edebfc1ed68532da6544120a624f8b143096d146d4b6c6

Observation 7eb2ced3-605c-44e7-922b-b64dde1fdb14 · inbound

ProWAFT: A ROMA-LPD Instance for Workload-Aware and Dynamic Fault Tolerance in FPGA-Based CNN Accelerators cites this paper.

ProWAFT: A ROMA-LPD Instance for Workload-Aware and Dynamic Fault Tolerance in FPGA-Based CNN Accelerators A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:33.659126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T15:23:48.748933Z digest=sha256:e327d9f6839f95b7be894f4abd62e1574e254349bdd4e6ee1dc9121ef2e9586e

Observation 9b31e906-c984-4c5c-9c27-529a9773e17b · inbound

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth cites this paper.

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth A Bounding Box is Worth One Token: Interleaving Layout and Text in a Large Language Model for Document Understanding

Reference 172

Resolution
unresolved
no resolver link, observed 2026-08-02T14:55:38.412643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:55:38.412643Z digest=sha256:796a76d1bf2b1f46b7c668083c2200bf9b32d4cf89bfecc89b635d8754a254b7