Pith. sign in

REVIEW

DTrOCR: Decoder-only Transformer for Optical Character Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2308.15996 v1 pith:X35JJ7E3 submitted 2023-08-30 cs.CV

classification cs.CV
keywords recognitiontextdecoder-onlydtrocrlanguagetransformercharactereffective
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Typical text recognition methods rely on an encoder-decoder structure, in which the encoder extracts features from an image, and the decoder produces recognized text from these features. In this study, we propose a simpler and more effective method for text recognition, known as the Decoder-only Transformer for Optical Character Recognition (DTrOCR). This method uses a decoder-only Transformer to take advantage of a generative language model that is pre-trained on a large corpus. We examined whether a generative language model that has been successful in natural language processing can also be effective for text recognition in computer vision. Our experiments demonstrated that DTrOCR outperforms current state-of-the-art methods by a large margin in the recognition of printed, handwritten, and scene text in both English and Chinese.

Discussion (0). Continue with ORCID to comment.

Pith tools