Pith. sign in

REVIEW 1 cited by

A Novel Pipeline for Improving Optical Character Recognition through Post-processing Using Natural Language Processing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.04245 v1 pith:MQL43BV5 submitted 2023-07-09 cs.CV cs.AI

classification cs.CVcs.AI
keywords applicationshandwrittenprintedaccuracycharactercharacterslanguagenatural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Optical Character Recognition (OCR) technology finds applications in digitizing books and unstructured documents, along with applications in other domains such as mobility statistics, law enforcement, traffic, security systems, etc. The state-of-the-art methods work well with the OCR with printed text on license plates, shop names, etc. However, applications such as printed textbooks and handwritten texts have limited accuracy with existing techniques. The reason may be attributed to similar-looking characters and variations in handwritten characters. Since these issues are challenging to address with OCR technologies exclusively, we propose a post-processing approach using Natural Language Processing (NLP) tools. This work presents an end-to-end pipeline that first performs OCR on the handwritten or printed text and then improves its accuracy using NLP.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Digitization of Document and Information Extraction using OCR

    cs.CV 2025-06 reject novelty 2.0 of 10

    The paper describes a standard OCR-plus-LLM extraction pipeline and reports unsupported accuracy percentages with no code, data, or baselines.

Pith tools