REVIEW 4 cited by
A Survey of Deep Learning Approaches for OCR and Document Understanding
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Documents are a core part of many businesses in many fields such as law, finance, and technology among others. Automatic understanding of documents such as invoices, contracts, and resumes is lucrative, opening up many new avenues of business. The fields of natural language processing and computer vision have seen tremendous progress through the development of deep learning such that these methods have started to become infused in contemporary document understanding systems. In this survey paper, we review different techniques for document understanding for documents written in English and consolidate methodologies present in literature to act as a jumping-off point for researchers exploring this area.
Forward citations
Cited by 4 Pith papers
-
The Documentation and Traceability Burden of the Indian EV Transition
The paper systematises India's EV compliance-document lifecycle into a two-layer evidence model, a six-stage lifecycle with four failure loci, an exergy-destruction analytic lens, and a six-problem research agenda.
-
Finding Needles in Images: Can Multimodal LLMs Locate Fine Details?
A new benchmark and method (Spot-IT) aim to improve multimodal LLMs' ability to locate fine details in documents, with reported significant gains.
-
Structured Data Extraction from Real Estate Documents using Clustering, Classification, and Large Language Models
A pipeline classifies 3965 real-estate questionnaires and extracts 35 structured attributes from 2781 selectable-text documents via DeepSeek R1, reporting Jaccard consistency 0.82.
-
Multi-Modal Vision vs. Text-Based Parsing: Benchmarking LLM Strategies for Invoice Processing
Across three invoice datasets, multimodal LLMs extract fields more accurately from raw images than from markdown converted by a parsing tool, with Gemini 2.5 Pro leading.
Discussion (0). Continue with ORCID to comment.