Pith. sign in

Paper Citation Record · LEDGER

SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2503.11576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.11576 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:26.990881Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:37:30.147692Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4f8d34df-db08-49f0-9a62-eb1f29194697 · inbound

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting cites this paper.

Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:26.990881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:42:26.990881Z digest=sha256:dea99732d87af1db588288cb9c1971b5a1f4b23f7dcc69542385a3665ccaa82f

Observation dc6b1429-fcc0-4f58-87ac-341cefcf2677 · inbound

PaddleOCR 3.0 Technical Report cites this paper.

PaddleOCR 3.0 Technical Report SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T23:24:28.327571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-14T23:24:28.094041Z digest=sha256:49d80a44871f175dd2ff52e03f58e57aa4b521675e71db05f15ba7c712e9f1ad

Observation 02a6f8ee-4043-4f86-b4ad-8467fcba4c50 · inbound

ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation cites this paper.

ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:10:18.495715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:10:18.495715Z digest=sha256:74feb298192bc724e2d106394adf9abcb59904d33c8c1d36845f0cb3043ee6ab

Observation 059b3d4f-e097-489e-adee-4c8cced1226e · inbound

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition cites this paper.

E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T10:52:13.940165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:52:13.940165Z digest=sha256:2c8930215be51f570057f48b6b6269d9d21bb21859ceda8824b1422ffa49f98b

Observation 0b8afe44-8a1a-46f0-bb21-6192ce1ac89a · inbound

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing cites this paper.

Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T14:20:25.170342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:20:25.170342Z digest=sha256:dabc8b63624a8e95591d263cfa6aa8cc4c5115bc789298c2f3aa68bc566752a8

Observation 2243c99a-0e45-49f1-8fff-ddf6d83300a6 · inbound

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing cites this paper.

MinerU2.5: A Decoupled Vision-Language Model for Efficient High-Resolution Document Parsing SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T13:25:32.028685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T13:25:31.884175Z digest=sha256:4b7dec1d0585982bfaa05e2feecb88d08b8ff39e9dd01a341d74f6c649efa5b2

Observation fd30d2fc-68a4-447b-be80-4c3bb7adff4b · inbound

DeepSeek-OCR: Contexts Optical Compression cites this paper.

DeepSeek-OCR: Contexts Optical Compression SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:35:50.751259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T04:35:50.647950Z digest=sha256:8929bb28fa84337260e7a58826261d1a720ff2d07c3a20fcf5a3b53442fb7e83

Observation a8e5ba99-f8e0-4699-9f45-d7b49523c9b2 · inbound

PubTables-v2: A new large-scale dataset for full-page and multi-page table extraction cites this paper.

PubTables-v2: A new large-scale dataset for full-page and multi-page table extraction SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T17:02:35.295207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:02:35.295207Z digest=sha256:121ab4848f41f87b6085513480aec298484d64ef799f784cc4fab0de9aa8680c

Observation 1a8d3d5e-c9d9-4741-bad0-dea1ffb02914 · inbound

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters cites this paper.

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T14:16:50.805641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:16:50.805641Z digest=sha256:c671ae1f97be1744a638ed080ae0d8ecfe5d09ce53164595ee7f4d3dbc212c6d

Observation 52ed26b1-3a17-4224-8537-68037e5b214d · inbound

DODO: Discrete OCR Diffusion Models cites this paper.

DODO: Discrete OCR Diffusion Models SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T22:27:42.627628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:27:42.627628Z digest=sha256:af6f218a16f25e8b780dc2b618409b73cc64dd9fde377e56826bb4eefe315246

Observation 914ea71a-7da4-48d3-9742-7e72075b1ed7 · inbound

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models cites this paper.

Q-Mask: Query-driven Causal Masks for Text Anchoring in OCR-Oriented Vision-Language Models SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:33:26.638872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T23:30:53.449935Z digest=sha256:39eaa357475b99505ecaee7cced020b578fcb6d3fd07705586cc22b1dbcfcaec

Observation 433cff05-eac4-4d32-b500-173a5f898e30 · inbound

DocAtlas: Multilingual Document Understanding Across 80+ Languages cites this paper.

DocAtlas: Multilingual Document Understanding Across 80+ Languages SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:02:58.627073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-14T21:02:45.148167Z digest=sha256:e998e54720ef44e3c9765a04a2089355538ac340a6b8ccbe15064b7019497a31

Observation 6db6664a-9a7e-45e2-93d5-55f7e1410f98 · inbound

TeleCom-Bench: How Far Are Large Language Models from Industrial Telecommunication Applications? cites this paper.

TeleCom-Bench: How Far Are Large Language Models from Industrial Telecommunication Applications? SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:28:14.663288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T11:23:34.256279Z digest=sha256:6467ca9797fcfbc4ed665ed4e5116db0db26d80c3507f132704951433523a2b2

Observation 8ddd3b6c-0e70-4645-914d-c65499a583f2 · inbound

Structured Layout Priors for Robust Out-of-Distribution Visual Document Understanding cites this paper.

Structured Layout Priors for Robust Out-of-Distribution Visual Document Understanding SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:05.101834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-20T05:48:34.771799Z digest=sha256:e04892efe5cf2c0685b678e97eaa2b9b50a291e335d49706a53ce6670e1849a5

Observation 2cf726a6-1fe4-4d53-b916-8553cd54bc25 · inbound

POTATR: A Lightweight Image-to-Graph Model for Page-Level Table Extraction cites this paper.

POTATR: A Lightweight Image-to-Graph Model for Page-Level Table Extraction SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:37:30.149616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T17:06:33.518620Z digest=sha256:6effd15fde62d2dfd44696dca1c199ce8a121593277001067508f68a12efc6b0