Pith. sign in

Paper Citation Record · LEDGER

UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2308.11592.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.11592 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:14:20.322963Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:54.680740Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b9623b53-7ac9-4dee-a312-fbe1bff6d34c · inbound

Cross-Modal Synergies: Unveiling the Potential of Motion-Aware Fusion Networks in Handling Dynamic and Static ReID Scenarios cites this paper.

Cross-Modal Synergies: Unveiling the Potential of Motion-Aware Fusion Networks in Handling Dynamic and Static ReID Scenarios UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T18:14:20.322963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:14:20.322963Z digest=sha256:89fc909f47480dfe4b2c65cd541f989355b9e3e6881b91d4784af9082e440975

Observation b518075c-ad1e-4f72-95b1-275264b0e729 · inbound

Trajectory Map-Matching in Urban Road Networks Based on RSS Measurements cites this paper.

Trajectory Map-Matching in Urban Road Networks Based on RSS Measurements UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-09T15:56:19.637295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T15:56:19.637295Z digest=sha256:0e71bdfc398f984a308e2ec829516d87fba0a926b8761def7548d3cf30fe3371

Observation 0e2b5b13-105c-4794-aadd-32dc8d11ce79 · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:25.553617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:25.553617Z digest=sha256:5184d26bcf747033cee40c815ddcdd9e767f03715e385fd73fa2ebc6e02697df

Observation 42e9637c-7a60-4daa-90e6-0a6a96d37501 · inbound

Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning cites this paper.

Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T15:29:17.222816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:29:17.222816Z digest=sha256:6d236aa14cbcc81b70526d2f32e186f61ea4f9e71a481c9f8abf23f125bfbbbf

Observation 491b1aae-a70a-43c4-8641-0e21201a8dae · inbound

SETransformer: A Hybrid Attention-Based Architecture for Robust Human Activity Recognition cites this paper.

SETransformer: A Hybrid Attention-Based Architecture for Robust Human Activity Recognition UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:20.539328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:20.539328Z digest=sha256:1c595aaa43a6c9819b3703c6baa1842f86da7b3fd7f09d70220ec62753c87457

Observation bfe7478d-1672-4ab9-b39b-a3d2e9c34a67 · inbound

Multimodal Tabular Reasoning with Privileged Structured Information cites this paper.

Multimodal Tabular Reasoning with Privileged Structured Information UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:58.839998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:58.839998Z digest=sha256:5f38eb76fcfda366299cef1ff4a9c4d6cd747a3927b3714eb4aa38a27c6aafd5

Observation d4ed33a4-f251-4927-860b-22ed8db00373 · inbound

DREAM: Document Reconstruction via End-to-end Autoregressive Model cites this paper.

DREAM: Document Reconstruction via End-to-end Autoregressive Model UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:45.203139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:45.203139Z digest=sha256:cb58da5d15739d3b1235af6c96867d9f94bb1abd46303da901a3ba2710891dac

Observation 9fc47ca3-2c32-4c31-ad41-99cd785f1d70 · inbound

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends cites this paper.

A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:42:04.332555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T04:38:49.512293Z digest=sha256:363253f0008bc70141d69eb1470f46fb8df894ef07537bc294a08d0a74202f3d

Observation 18f6e641-c3e9-4bdb-8b07-e38c0bab17d3 · inbound

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data cites this paper.

LPCAN: Lightweight Pyramid Cross-Attention Network for Rail Surface Defect Detection Using RGB-D Data UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:50.176177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:50.176177Z digest=sha256:37e3f34c7e266835b15705e28a60f96a3c5dac88281265fbc51211381db7592d

Observation 66d31a73-531b-4a3a-b04b-5cc2dc5084f7 · inbound

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method cites this paper.

Knowledge-Embedded and Hypernetwork-Guided Few-Shot Substation Meter Defect Image Generation Method UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T10:43:42.195991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:43:42.195991Z digest=sha256:7c8c3b0eeb785ff846c859e11fbdc7d79f331f7641c06c4e84d047d6f68564ff

Observation 65d1ca06-3817-414e-9a67-08313bfafc78 · inbound

Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction cites this paper.

Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:03:12.272128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T20:01:57.029143Z digest=sha256:b9108ddac7851203e11e9a354f676b1d731831150dd9e707ac9857adbced4d78

Observation 3ea70519-a7a2-4a26-b304-bb5115688556 · inbound

Feature Perturbation Pool-based Fusion Network for Unified Multi-Class Industrial Defect Detection cites this paper.

Feature Perturbation Pool-based Fusion Network for Unified Multi-Class Industrial Defect Detection UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:29:21.633982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T03:28:11.747826Z digest=sha256:7af4bf330b1cb08b9f10b43dae08a37f3486c0200ffbe3bcde1fb75846c70ace

Observation 40e68cff-df95-4c4b-a0cb-0a2a6d77cd14 · inbound

Lightweight Real-Time Rendering Parameter Optimization via XGBoost-Driven Lookup Tables cites this paper.

Lightweight Real-Time Rendering Parameter Optimization via XGBoost-Driven Lookup Tables UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:21:14.845234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T17:25:18.561710Z digest=sha256:7e83711461674d766c991e644c9260ebead9fc9d7f66fae894cc774730ce3a08

Observation 173571e1-59a4-4d9d-bcfd-0229e46e3375 · inbound

Image Classification via Random Dilated Convolution with Multi-Branch Feature Extraction and Context Excitation cites this paper.

Image Classification via Random Dilated Convolution with Multi-Branch Feature Extraction and Context Excitation UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:22:02.896411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T17:07:55.177114Z digest=sha256:caa52e58c1edaf89f3054df66d08bde7664976a74818b597e458b8bbaa266f90

Observation 74eb008a-5670-42f5-a095-8928d9f4af5d · inbound

Gait Recognition via Deep Residual Networks and Multi-Branch Feature Fusion cites this paper.

Gait Recognition via Deep Residual Networks and Multi-Branch Feature Fusion UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:01:27.637705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T08:33:09.327681Z digest=sha256:87193664de0eba6deccd691c0da0d9b24e0eebdc2b12f6531433ca628b1df80c

Observation d5f1614e-d80a-4f87-a011-61fffda2d4b3 · inbound

Multi-Branch Non-Homogeneous Image Dehazing via Concentration Partitioning and Image Fusion cites this paper.

Multi-Branch Non-Homogeneous Image Dehazing via Concentration Partitioning and Image Fusion UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:08.682659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T20:03:36.551943Z digest=sha256:f9be2d68383cb6cb07d3c27249d9a1289ca9608a23a3a09a3e4b526ab88c3b57

Observation a254b20e-b1bc-468f-9d04-297da62be12c · inbound

Local Spatiotemporal Convolutional Network for Robust Gait Recognition cites this paper.

Local Spatiotemporal Convolutional Network for Robust Gait Recognition UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:08:29.320024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:07:25.917495Z digest=sha256:751a6db7ffc27809a7a9573b056c027793212bcdbb98ad2e8cfc2db4a7ce78c7

Observation 9921cc2d-3f12-4dae-8d28-12f01bd57dd3 · inbound

Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting cites this paper.

Do You Need Text Rectification? Soft Attention Mask Embedding for Rectification-Free Scene Text Spotting UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:33:14.584828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T11:28:51.711892Z digest=sha256:73664c682d90ae763872fe12870a46d3a6ac7a050d091cf09e1735363613ed20

Observation e39545d6-7c74-40c9-9b46-2e5e934fec8a · inbound

An LMM for Precisely Grounding Elements in Documents cites this paper.

An LMM for Precisely Grounding Elements in Documents UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.682949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T01:51:50.043216Z digest=sha256:71bae17a426f86095f9177555237cd5768c05c628390f8afbfec2f02e8889bb9

Observation 01858738-d7f0-413c-ae03-36fc32022231 · inbound

DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning cites this paper.

DAIN: Dynamic Agent-Based Interaction Network for Efficient and Collaborative Multimodal Reasoning UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:04:21.403127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T06:01:20.803078Z digest=sha256:477aedb80f145bfa07cc1efb954d4afe49fe129df474db04bb584d4930d004f6

Observation fb14e358-776c-4cd9-92c6-922475fe23a9 · inbound

ProWAFT: A ROMA-LPD Instance for Workload-Aware and Dynamic Fault Tolerance in FPGA-Based CNN Accelerators cites this paper.

ProWAFT: A ROMA-LPD Instance for Workload-Aware and Dynamic Fault Tolerance in FPGA-Based CNN Accelerators UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:28:33.653477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-03T15:23:48.748933Z digest=sha256:c0937fe38d7357574ba1a906412b987b6c7e341c3f443da3e08023027428192d

Observation 9db67247-d520-4e54-8ab4-e626c6140e15 · inbound

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth cites this paper.

DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth UniDoc: A Universal Large Multimodal Model for Simultaneous Text Detection, Recognition, Spotting and Understanding

Reference 217

Resolution
unresolved
no resolver link, observed 2026-08-02T14:55:42.466530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:55:42.466530Z digest=sha256:c26c0a1456fac877d96303a0f0d7a92016bc07ae6738593f4bc9c71f26cb04df