Pith. sign in

Paper Citation Record · LEDGER

UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2308.11592.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.11592 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:25.553617Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:54.680740Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0e2b5b13-105c-4794-aadd-32dc8d11ce79 · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:25.553617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:25.553617Z digest=sha256:5184d26bcf747033cee40c815ddcdd9e767f03715e385fd73fa2ebc6e02697df

Observation 42e9637c-7a60-4daa-90e6-0a6a96d37501 · inbound

Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning cites this paper.

Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:17.222816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:17.222816Z digest=sha256:6d236aa14cbcc81b70526d2f32e186f61ea4f9e71a481c9f8abf23f125bfbbbf

Observation 491b1aae-a70a-43c4-8641-0e21201a8dae · inbound

SETransformer: A Hybrid Attention-Based Architecture for Robust Human Activity Recognition cites this paper.

SETransformer: A Hybrid Attention-Based Architecture for Robust Human Activity Recognition UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:20.539328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:20.539328Z digest=sha256:1c595aaa43a6c9819b3703c6baa1842f86da7b3fd7f09d70220ec62753c87457

Observation bfe7478d-1672-4ab9-b39b-a3d2e9c34a67 · inbound

Multimodal Tabular Reasoning with Privileged Structured Information cites this paper.

Multimodal Tabular Reasoning with Privileged Structured Information UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:58.839998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:58.839998Z digest=sha256:5f38eb76fcfda366299cef1ff4a9c4d6cd747a3927b3714eb4aa38a27c6aafd5

Observation d4ed33a4-f251-4927-860b-22ed8db00373 · inbound

DREAM: Document Reconstruction via End-to-end Autoregressive Model cites this paper.

DREAM: Document Reconstruction via End-to-end Autoregressive Model UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.203139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.203139Z digest=sha256:cb58da5d15739d3b1235af6c96867d9f94bb1abd46303da901a3ba2710891dac

Observation 9fc47ca3-2c32-4c31-ad41-99cd785f1d70 · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.332555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:10476e0d0c53828c8453ec11104a50b9c40527074f438b13a2068100a5ff1133

Observation 18f6e641-c3e9-4bdb-8b07-e38c0bab17d3 · inbound

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data cites this paper.

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:50.176177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:50.176177Z digest=sha256:37e3f34c7e266835b15705e28a60f96a3c5dac88281265fbc51211381db7592d

Observation 66d31a73-531b-4a3a-b04b-5cc2dc5084f7 · inbound

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method cites this paper.

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:42.195991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:42.195991Z digest=sha256:a804c7c3a789b0d28796ad1cd304e79e928afdf9b2d1e44f170a2216c2c1d5a9

Observation 65d1ca06-3817-414e-9a67-08313bfafc78 · inbound

Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction cites this paper.

Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:03:12.272128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T20:01:57.029143Z digest=sha256:ab5dccb8f25e74ec245111cd0df391b22aaae15aae46dc7666b9eac0744a4bd3

Observation 3ea70519-a7a2-4a26-b304-bb5115688556 · inbound

Feature Perturbation Pool-based Fusion Network for Unified Multi-Class Industrial Defect Detection cites this paper.

Feature Perturbation Pool-based Fusion Network for Unified Multi-Class Industrial Defect Detection UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:29:21.633982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T03:28:11.747826Z digest=sha256:9203804ce5916d77778d759003f0278198a4f280ad23ec1e55e73a262c846c94

Observation 40e68cff-df95-4c4b-a0cb-0a2a6d77cd14 · inbound

Lightweight Real-Time Rendering Parameter Optimization via XGBoost-Driven Lookup Tables cites this paper.

Lightweight Real-Time Rendering Parameter Optimization via XGBoost-Driven Lookup Tables UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:21:14.845234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T17:25:18.561710Z digest=sha256:075fcecb0c6a9db9d4e8eb2790ded12024eb8b484becc9a62fbb0d0ce5f73afc

Observation 173571e1-59a4-4d9d-bcfd-0229e46e3375 · inbound

Image Classification via Random Dilated Convolution with Multi-Branch Feature Extraction and Context Excitation cites this paper.

Image Classification via Random Dilated Convolution with Multi-Branch Feature Extraction and Context Excitation UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:22:02.896411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T17:07:55.177114Z digest=sha256:4cdad4c027c425f64ef10aa98366a24bfa187ac656c8b23c718f69ec0b6fb6e5

Observation 74eb008a-5670-42f5-a095-8928d9f4af5d · inbound

Gait Recognition via Deep Residual Networks and Multi-Branch Feature Fusion cites this paper.

Gait Recognition via Deep Residual Networks and Multi-Branch Feature Fusion UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:01:27.637705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T08:33:09.327681Z digest=sha256:13c9081b2a13cbdc3944b86bf786d8d030cdaee9b0f4db6813672645f6f129b4

Observation d5f1614e-d80a-4f87-a011-61fffda2d4b3 · inbound

Multi-Branch Non-Homogeneous Image Dehazing via Concentration Partitioning and Image Fusion cites this paper.

Multi-Branch Non-Homogeneous Image Dehazing via Concentration Partitioning and Image Fusion UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:08.682659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T20:03:36.551943Z digest=sha256:56fdb6edaf06179e4ec28486f535ee724767376e217a36ad55d6d983445f9347

Observation a254b20e-b1bc-468f-9d04-297da62be12c · inbound

Local Spatiotemporal Convolutional Network for Robust Gait Recognition cites this paper.

Local Spatiotemporal Convolutional Network for Robust Gait Recognition UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:08:29.320024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:07:25.917495Z digest=sha256:98da994651a974e8b9ef3cb8abd44decb214259909118594b860e5c4e40ca34c

Observation 9921cc2d-3f12-4dae-8d28-12f01bd57dd3 · inbound

Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting cites this paper.

Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:33:14.584828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T11:28:51.711892Z digest=sha256:3d748e5cfb5446c0beb11df94a453ac173f9008b66afe6b96a993039f96d1dfe

Observation e39545d6-7c74-40c9-9b46-2e5e934fec8a · inbound

An LMM for Precisely Grounding Elements in Documents cites this paper.

An LMM for Precisely Grounding Elements in Documents UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.682949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T01:51:50.043216Z digest=sha256:08aa048e6f41e0d57f6813e7f1af1d6feb8609958311a26d289039473c9f1c12

Observation 01858738-d7f0-413c-ae03-36fc32022231 · inbound

DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning cites this paper.

DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:04:21.403127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:01:20.803078Z digest=sha256:0fd42b6e21d93b68eaa2d0b7ebb71628ffe11f84c5e676941f1277298c9d678d

Observation fb14e358-776c-4cd9-92c6-922475fe23a9 · inbound

ProWAFT: A ROMA-LPD Instance for Workload-Aware and Dynamic Fault Tolerance in FPGA-Based CNN Accelerators cites this paper.

ProWAFT: A ROMA-LPD Instance for Workload-Aware and Dynamic Fault Tolerance in FPGA-Based CNN Accelerators UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:33.653477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T15:23:48.748933Z digest=sha256:60d6602f0ee6c021f8fc40d60bad9434709b6007f6b79a456d2a0b604fd55d0c

Observation 9db67247-d520-4e54-8ab4-e626c6140e15 · inbound

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth cites this paper.

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 217

Resolution
unresolved
no resolver link, observed 2026-08-02T14:55:42.466530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:55:42.466530Z digest=sha256:c26c0a1456fac877d96303a0f0d7a92016bc07ae6738593f4bc9c71f26cb04df