Pith. sign in

REVIEW 7 cited by

SmolDocling: An ultra-compact vision-language model for end-to-end multi-modal document conversion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.11576 v1 pith:JJCLIWZU submitted 2025-03-14 cs.CV

classification cs.CV
keywords documentmodelsmoldoclingconversionend-to-endmodelsvision-languageavailable
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce SmolDocling, an ultra-compact vision-language model targeting end-to-end document conversion. Our model comprehensively processes entire pages by generating DocTags, a new universal markup format that captures all page elements in their full context with location. Unlike existing approaches that rely on large foundational models, or ensemble solutions that rely on handcrafted pipelines of multiple specialized models, SmolDocling offers an end-to-end conversion for accurately capturing content, structure and spatial location of document elements in a 256M parameters vision-language model. SmolDocling exhibits robust performance in correctly reproducing document features such as code listings, tables, equations, charts, lists, and more across a diverse range of document types including business documents, academic papers, technical reports, patents, and forms -- significantly extending beyond the commonly observed focus on scientific papers. Additionally, we contribute novel publicly sourced datasets for charts, tables, equations, and code recognition. Experimental results demonstrate that SmolDocling competes with other Vision Language Models that are up to 27 times larger in size, while reducing computational requirements substantially. The model is currently available, datasets will be publicly available soon.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PubTables-v2: A new large-scale dataset for full-page and multi-page table extraction

    cs.CV 2025-12 conditional novelty 7.0 of 10

    PubTables-v2 is a large annotated dataset for table extraction spanning cropped tables, full pages, and full documents, including the first large benchmark of multi-page tables.

  2. DODO: Discrete OCR Diffusion Models

    cs.CV 2026-02 conditional novelty 6.0 of 10

    Block-based discrete diffusion can transcribe documents in parallel, roughly matching autoregressive OCR accuracy while cutting inference time by up to about 3x in a lower-accuracy fast variant.

  3. ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation

    cs.HC 2025-05 conditional novelty 6.0 of 10

    A fully automated pipeline generates a 222.5K-pair synthetic chart dataset with 27 chart types and 11 plotting libraries, and a GPT-4o-judged benchmark shows current open-weights VLMs still underperform on chart-to-co...

  4. Dolphin: Document Image Parsing via Heterogeneous Anchor Prompting

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A two-stage model that analyzes page layout first, then parses text, tables, and formulas in parallel, reports state-of-the-art accuracy and speed on the benchmarks it evaluates.

  5. UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters

    cs.CV 2025-12 conditional novelty 5.0 of 10

    A 0.1B-parameter text/formula recognition model trained on a new 40M-sample dataset matches or beats much larger OCR models and runs 2-9× faster.

  6. E-ARMOR: Edge case Assessment and Review of Multilingual Optical Character Recognition

    cs.CL 2025-09 reject novelty 4.0 of 10

    Their custom PaddleOCR-based system achieves the best F1 (0.46), fastest latency (0.17 s/image), and lowest cost ($0.006/1k images) among seven OCR systems on a private 54-language benchmark.

  7. Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing

    cs.CL 2025-08 conditional novelty 4.0 of 10

    Across three invoice datasets, multimodal LLMs extract fields more accurately from raw images than from markdown converted by a parsing tool, with Gemini 2.5 Pro leading.

Pith tools