REVIEW 6 cited by
TrOCR: Transformer-based Optical Character Recognition with Pre-trained Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Text recognition is a long-standing research problem for document digitalization. Existing approaches are usually built based on CNN for image understanding and RNN for char-level text generation. In addition, another language model is usually needed to improve the overall accuracy as a post-processing step. In this paper, we propose an end-to-end text recognition approach with pre-trained image Transformer and text Transformer models, namely TrOCR, which leverages the Transformer architecture for both image understanding and wordpiece-level text generation. The TrOCR model is simple but effective, and can be pre-trained with large-scale synthetic data and fine-tuned with human-labeled datasets. Experiments show that the TrOCR model outperforms the current state-of-the-art models on the printed, handwritten and scene text recognition tasks. The TrOCR models and code are publicly available at \url{https://aka.ms/trocr}.
Forward citations
Cited by 6 Pith papers
-
Predicting the Past: Estimating Historical Appraisals with OCR and Machine Learning
A hand-annotated dataset of 1933 Hamilton County property appraisals is extracted with template-aligned OCR, and a random forest trained on contemporary features estimates those historical values with about 17% MAPE.
-
HAND: Hierarchical Attention Network for Multi-Scale Handwritten Document Recognition and Layout Analysis
HAND is an end-to-end handwritten document recognition network that jointly performs text recognition and layout analysis and reports state-of-the-art results on READ 2016, scaling to triple-page documents.
-
WriteViT: Handwritten Text Generation with Vision Transformer
A ViT-based GAN framework for one-shot handwriting synthesis that reports state-of-the-art FID/KID on IAM and VNOnDB and improves HTR performance when used as a data augmenter.
-
Arabic-Nougat: Fine-Tuning Vision Transformers for Arabic OCR and Markdown Extraction
Arabic-specific fine-tunes of Nougat, trained on a synthetic Hindawi-derived dataset, report high in-distribution BLEU and structure accuracy for Arabic book page OCR.
-
LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR
A 1B end-to-end VLM reports 83.2 on OlmOCR-Bench after excluding the headers/footers category, plus a new image-localization benchmark.
-
Low-Resource Language Processing: An OCR-Driven Summarization and Translation Pipeline
A pipeline combining Tesseract, Gemini, and Google Translate claims 88% classification accuracy and moderate BLEU/ROUGE scores, but provides no reproducible evaluation.
Discussion (0). Continue with ORCID to comment.