Pith. sign in

Paper Citation Record · LEDGER

DocFusion: A Unified Framework for Document Parsing Tasks

As of 15 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2412.12505.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12505 v2

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:04:51.807410Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:25.216463Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:42:32.291738Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved19
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f9219257-a460-45f9-8cf5-633234105552 · outbound

This paper cites Nougat: Neural Optical Understanding for Academic Documents.

DocFusion: A Unified Framework for Document Parsing Tasks Nougat: Neural Optical Understanding for Academic Documents

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.712943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.712943Z digest=sha256:45bf0e7d6a5e4c9f052c914cde69d347712504268de3fe3f18af052bda81cce7

Observation 7735b225-1c7f-43af-a4ab-9a9bf45db3dc · outbound

This paper cites DaViT: Dual Attention Vision Transformers.

DocFusion: A Unified Framework for Document Parsing Tasks DaViT: Dual Attention Vision Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.729549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.729549Z digest=sha256:b7f4f0a1ee4e721ae456b5835d40b7b1affafd207bfb10016126aafa6bc9351f

Observation f0e9cf88-61f0-4e5c-aec6-84e7e3da4ffd · outbound

This paper cites LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking.

DocFusion: A Unified Framework for Document Parsing Tasks LayoutLMv3: Pre-training for Document AI with Unified Text and Image Masking

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.738129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.738129Z digest=sha256:ef6e622916145a8d0ca5311521e42bc7d29ce3229fac2a4d4cf2ff4831459915

Observation ddb24156-520e-48da-87e5-5a8b84046204 · outbound

This paper cites YOLOv11: An Overview of the Key Architectural Enhancements.

DocFusion: A Unified Framework for Document Parsing Tasks YOLOv11: An Overview of the Key Architectural Enhancements

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.742600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.742600Z digest=sha256:2376460357fb1d6f4f60322947e4f98481b2ca8d90409ba632191012a5db8c31

Observation 07c475fb-f1fd-4ec2-835c-2a6c7b53af61 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

DocFusion: A Unified Framework for Document Parsing Tasks LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.746770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.746770Z digest=sha256:9541667edb0280d255cb86e32c68f458ccb3bd48f82fb6da97be620ff8ff11ed

Observation d95ab6b9-88fb-4454-992b-9a894289b9da · outbound

This paper cites TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document.

DocFusion: A Unified Framework for Document Parsing Tasks TextMonkey: An OCR-Free Large Multimodal Model for Understanding Document

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.751083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.751083Z digest=sha256:2fdfb17c284d66557982c9e69ebb57ed37e88bac3c1de37506fbd30aaa778263

Observation 6672cdf6-4fef-4255-9a2b-d12bd3b4d655 · outbound

This paper cites Preprint, arXiv:2406.17148.

DocFusion: A Unified Framework for Document Parsing Tasks Preprint, arXiv:2406.17148

Reference 13

Resolution
verified exact
raw_fallback, observed 2026-08-11T14:04:52.069796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:04:51.754902Z digest=sha256:6e166b24e7ebd8f38ad36d8a3e0ba0a1e4d3c1258ee0fa75a271026d47245e7f

Observation f14083fb-eb4d-4f45-984a-5e3af3918694 · outbound

This paper cites Visually Guided Generative Text-Layout Pre-training for Document Intelligence.

DocFusion: A Unified Framework for Document Parsing Tasks Visually Guided Generative Text-Layout Pre-training for Document Intelligence

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.758577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.758577Z digest=sha256:409fe5dfa9261b1141ad1545335711650d04cab1f365f7f964de66e03bcd66bd

Observation b29c159f-46eb-4936-809a-6c78bc6dacd7 · outbound

This paper cites Accessed: 2024-02-29, cited in pages 1, 2, 4, 6,.

DocFusion: A Unified Framework for Document Parsing Tasks Accessed: 2024-02-29, cited in pages 1, 2, 4, 6,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:52.190983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:04:51.762483Z digest=sha256:ff0661ad31c56b32bc1d5a9ba6f58e3d41137053e687cee3ef0cab9201503b06

Observation be84449a-d9f5-4460-b4f2-16ba2cfc86a7 · outbound

This paper cites RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-ranking.

DocFusion: A Unified Framework for Document Parsing Tasks RocketQAv2: A Joint Training Method for Dense Passage Retrieval and Passage Re-ranking

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.769933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.769933Z digest=sha256:465fdf60fe3fb67180d88708953c166ec057f61beed68a83d2f75fbb9bebf5c0

Observation 3a165f9f-8af0-404f-9027-4c9205e448a9 · outbound

This paper cites General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model.

DocFusion: A Unified Framework for Document Parsing Tasks General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.777751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.777751Z digest=sha256:a6171a3d47443f3f7c17f114b405b33cf949de77e05310b47ee9c856419761ab

Observation c52db8e3-b3d1-4815-8219-9ecf91af8376 · outbound

This paper cites arXiv preprint arXiv:2406.11633.

DocFusion: A Unified Framework for Document Parsing Tasks arXiv preprint arXiv:2406.11633

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.785069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.785069Z digest=sha256:6ae12546f92de60b35e77b7d1177995ebb6bc5ef4fc7a2e05ec030468f10c1a0

Observation 114ebe78-1c71-4478-bba7-2c83e5ff44dc · outbound

This paper cites Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks.

DocFusion: A Unified Framework for Document Parsing Tasks Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.788539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.788539Z digest=sha256:6a7457b14f52d8420f27c3b58ea1f2e944138029253e156de1d2d993e0fe2508

Observation eba26595-eb38-46da-afe9-ed752e090211 · outbound

This paper cites LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding.

DocFusion: A Unified Framework for Document Parsing Tasks LayoutLMv2: Multi-modal Pre-training for Visually-Rich Document Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.792406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.792406Z digest=sha256:ce356552fa15cec7e3011581cbc9c7be8dc9732955218a014b5f642c53848930

Observation eea32079-4f42-49f9-8e56-720633a0d8b3 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

DocFusion: A Unified Framework for Document Parsing Tasks UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.796087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.796087Z digest=sha256:36976577b895a626b8f9694ef361f0d8e243b4f720b0af7b15ef1b678489964c

Observation 2ec46c1f-b940-4f02-a4a2-341146ca7eb1 · outbound

This paper cites Syntax-Aware Network for Handwritten Mathematical Expression Recognition.

DocFusion: A Unified Framework for Document Parsing Tasks Syntax-Aware Network for Handwritten Mathematical Expression Recognition

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.799769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.799769Z digest=sha256:38a223d654c868add49f348d9a176425f357a2d9dc7557d7553f6e398d9fbf06

Observation 3930debd-acc2-46a0-9c65-71547acf2492 · outbound

This paper cites Adversarial Retriever-Ranker for dense text retrieval.

DocFusion: A Unified Framework for Document Parsing Tasks Adversarial Retriever-Ranker for dense text retrieval

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.803689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.803689Z digest=sha256:78495f7b9443124a19d8d454f9c7e57a323d774bb733d666bb439d2f1e179f54

Observation c693a7b4-4e5e-4b55-b153-e28abb9ce2f4 · outbound

This paper cites However, Latex has not been the mainstream approach for TR tasks in recent times, resulting in a limited number of TR models available for comparison in the main experiment.

DocFusion: A Unified Framework for Document Parsing Tasks However, Latex has not been the mainstream approach for TR tasks in recent times, resulting in a limited number of TR models available for comparison in the main experiment

Reference 27

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T14:04:52.158585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:04:51.807410Z digest=sha256:728e663fe5a0cc71a66267944156768743f08be7a59e3f0626e770fc028dd491

Observation ca7fc0be-3469-4690-afdd-b1d369141bcf · outbound

This paper cites YOLOv10: Real-Time End-to-End Object Detection.

DocFusion: A Unified Framework for Document Parsing Tasks YOLOv10: Real-Time End-to-End Object Detection

Reference 2007

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.773913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.773913Z digest=sha256:2760f5ba1eea35d659813a9c438e4e850e543172afd46fc8800f66bae97d5721

Observation ec6acdb3-6e39-4036-88a0-2c5cd12af7cf · outbound

This paper cites In 2017 IEEE International Conference on Computer Vision (ICCV), pages 2223–2231.

DocFusion: A Unified Framework for Document Parsing Tasks In 2017 IEEE International Conference on Computer Vision (ICCV), pages 2223–2231

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:52.211337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:04:51.717239Z digest=sha256:4e650b33c2050226334fcc5234e108cb9d0ac9dc7d6255011424acfeef8e11c1

Observation 0d2cea3a-35c7-470b-8072-70fc008ad379 · outbound

This paper cites In 2018 13th IAPR International Workshop on Document Analysis Systems (DAS).

DocFusion: A Unified Framework for Document Parsing Tasks In 2018 13th IAPR International Workshop on Document Analysis Systems (DAS)

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:52.169695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:04:51.781419Z digest=sha256:0d2d66c02c856c968f9339f8b183f47809813af6a5009d47a20af483c620982a

Observation 0533ed28-c0ca-424c-ae7a-c7841daa0374 · outbound

This paper cites In 2019 International Conference on Document Analysis and Recognition (ICDAR), pages 142–147.

DocFusion: A Unified Framework for Document Parsing Tasks In 2019 International Conference on Document Analysis and Recognition (ICDAR), pages 142–147

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:52.179841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:04:51.766461Z digest=sha256:f7ed696fdc4fed42e3bd1ae89b910e7b97dfa9b4341a1666cfeb0326e39acdf6

Observation f6f59ad0-6610-4672-af1b-dc68a737765c · outbound

This paper cites End-to-End Object Detection with Transformers.

DocFusion: A Unified Framework for Document Parsing Tasks End-to-End Object Detection with Transformers

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.720965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.720965Z digest=sha256:b742df741c6f7cd32c4f6b5424d89a2f27cdf0bc50e5db595992b02a1b55c30a

Observation 6054a1b2-20e5-4d23-b4e1-69d19817b9ca · outbound

This paper cites In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).

DocFusion: A Unified Framework for Document Parsing Tasks In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:52.200948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:04:51.725583Z digest=sha256:198933c9bd996c13833ed845e1689aef6c698c970a4f7577028104286d4a0f07

Observation 90205521-6102-4b43-93cc-802a336fa179 · outbound

This paper cites Accessed: 2024-02-29, cited in pages 1, 2, 3, 7, 10,.

DocFusion: A Unified Framework for Document Parsing Tasks Accessed: 2024-02-29, cited in pages 1, 2, 3, 7, 10,

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:04:52.222424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T14:04:51.709022Z digest=sha256:85f0d105efbc1ce72188e497552ef090a6a7985352ffb81d702635a9d61d6d76

Observation 6c35c1eb-d621-4119-963a-855ec8ee9c95 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

DocFusion: A Unified Framework for Document Parsing Tasks Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.704202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.704202Z digest=sha256:69259b4c7a0f1dad3827933a31866dd113d511e302b8bd79f1cfdf25598f5094

Observation 335f972b-dc07-42f6-8aa4-5404dc73ac78 · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

DocFusion: A Unified Framework for Document Parsing Tasks Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T14:04:51.733707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:04:51.733707Z digest=sha256:fee1dfb0e6c5c1e7e99988441e242c6a62eb802b123231caad5b30c0775a3478

Pith citing papers

Observation d0af0262-56e8-4e72-8c9b-7a894f0d0fd4 · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting DocFusion: A Unified Framework for Document Parsing Tasks

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:42:32.401854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T15:42:25.216463Z digest=sha256:f2ab786a7a0162a03de7836f66d16842b08052e22e897b235ff353dd38d98446