Pith. sign in

REVIEW 4 cited by

A Survey of Deep Learning Approaches for OCR and Document Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.13534 v2 pith:Y7MSZMCS submitted 2020-11-27 cs.CL cs.CVcs.IRcs.LG

classification cs.CLcs.CVcs.IRcs.LG
keywords understandingdocumentdocumentsmanydeepfieldslearningsurvey
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Documents are a core part of many businesses in many fields such as law, finance, and technology among others. Automatic understanding of documents such as invoices, contracts, and resumes is lucrative, opening up many new avenues of business. The fields of natural language processing and computer vision have seen tremendous progress through the development of deep learning such that these methods have started to become infused in contemporary document understanding systems. In this survey paper, we review different techniques for document understanding for documents written in English and consolidate methodologies present in literature to act as a jumping-off point for researchers exploring this area.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The Documentation and Traceability Burden of the Indian EV Transition

    cs.CE 2026-07 conditional novelty 7.0 of 10

    The paper systematises India's EV compliance-document lifecycle into a two-layer evidence model, a six-stage lifecycle with four failure loci, an exergy-destruction analytic lens, and a six-problem research agenda.

  2. Finding Needles in Images: Can Multimodal LLMs Locate Fine Details?

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    A new benchmark and method (Spot-IT) aim to improve multimodal LLMs' ability to locate fine details in documents, with reported significant gains.

  3. Structured Data Extraction from Real Estate Documents using Clustering, Classification, and Large Language Models

    cs.CV 2026-07 unverdicted novelty 4.0 of 10

    A pipeline classifies 3965 real-estate questionnaires and extracts 35 structured attributes from 2781 selectable-text documents via DeepSeek R1, reporting Jaccard consistency 0.82.

  4. Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing

    cs.CL 2025-08 conditional novelty 4.0 of 10

    Across three invoice datasets, multimodal LLMs extract fields more accurately from raw images than from markdown converted by a parsing tool, with Gemini 2.5 Pro leading.

Pith tools