Pith. sign in

REVIEW 3 major objections 5 minor 1 cited by

Handwritten Text Recognition: A Survey

T0 review · 3 major / 5 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read This survey organizes handwritten text recognition by reading-order complexity and reports that beyond-line methods cluster tightly on IAM, with line-adaptation approaches leading.

desk verdict A useful reading-order taxonomy for HTR surveys; the benchmark comparison is the weak joint and should be fixed before publication. read the letter →

arxiv 2502.08417 v1 pith:SWB3UVDL submitted 2025-02-12 cs.CV cs.AI

classification cs.CVcs.AI
keywords handwrittentextrecognitionreadingorderline-leveltranscriptiondocument-levelconnectionisttemporalclassificationsegmentation-freebenchmarkingsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey argues that the field of Handwritten Text Recognition is best organized by the number of reading-order directions a model must follow. It divides methods into those that transcribe up to one line (words and lines, a single reading direction) and those that go beyond the line (paragraphs and documents, where a second or third direction appears). Within that framework it compares published results on the IAM benchmark and reports that beyond-line error rates cluster tightly, with line-adaptation approaches (masking and unfolding) ahead of unconstrained document-level models. The takeaway is that for paragraph-level IAM, the decisive step is still converting the page into line-level units and then using established line transcription techniques. A reader should care because this gives a simple axis for comparing a fragmented literature and a concrete claim about where the current practical bottleneck is.

What carries the argument

The central object is the reading-order hierarchy (Definitions 1 through 3) together with the image-to-sequence collapse function $r(\cdot)$, which maps a 2D feature map to a 1D sequence. Up-to-line methods rely on $r(\cdot)$ collapsing vertical features into frames; beyond-line methods either learn a mask that selects lines before collapsing (attention masking), reshape the feature map to unfold lines into one long sequence (line unfolding), or bypass the collapse entirely with a Transformer decoder that learns an arbitrary reading order (unconstrained). The survey uses this machinery to classify every method and to explain why the vertical collapse is the critical technical barrier at paragraph and document level.

What would settle it

Re-run the methods compared in Tables II and III under a single protocol: identical IAM splits, identical character sets, identical language-model settings, and identical synthetic pre-training. If an unconstrained model such as DAN or FPHR then matches or beats the masking and unfolding models on a document-level corpus, the survey's conclusion that line-level adaptation is the key factor would be overturned.

Watch

Extended reading notes

Core claim

The paper's central discovery is taxonomic: the meaningful complexity boundary in HTR is not image size but reading order. It defines line-level HTR as input with one reading direction, paragraph-level as two directions (line direction plus line-to-line direction), and document-level as an arbitrary third direction. It then maps the entire methodological literature onto this axis: up-to-line methods split into handcrafted pipelines (explicit or implicit segmentation) and end-to-end models (CTC, sequence-to-sequence, hybrid); beyond-line methods split into attention masking, line unfolding, and unconstrained approaches. On the empirical side, the survey collects IAM character error rates and finds that beyond-line systems are tightly clustered, with the best masking and unfolding systems (VAN and Origaminet) reaching about 4.6 to 4.7 percent CER, while the best unconstrained system (MSDocTr-Lite) reports 6.4 percent. It concludes that the key factor lies in adapting the document into a line-level structure for transcription using traditional methods.

Load-bearing premise

The survey's central performance conclusions assume that the CER values collected from different papers are directly comparable, even though the original works use different dataset splits, character sets, language models, and synthetic pre-training; the paper itself acknowledges this lack of a unified comparison framework.

Editorial extensions

If this is right

  • If the taxonomy is right, future method papers should state which reading-order level they target, since the comparison axis is not image size but the number of reading directions.
  • If the performance conclusion is right, practitioners building paragraph-level HTR on IAM-like data should prefer masking or unfolding plus CTC over unconstrained document models, because they are simpler and match or beat them.
  • If the tight clustering is real, IAM no longer discriminates among beyond-line approaches; evaluations should shift to datasets like Rimes or Bozen that exercise complex layouts.
  • If the framework holds, benchmarking should standardize tokenization and synthetic-data reporting, because those choices currently change the meaning of the reported character error rate.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: the same reading-order axis could organize related tasks like layout analysis and document understanding, where reading order is usually treated as a separate post-processing module.
  • Inference: if the line-adaptation result generalizes, then synthetic pre-training for document-level HTR should be designed to produce line-adaptable representations, not only full-page images.
  • Inference: the taxonomy predicts that a document-level model that can switch reading orders without retraining would be a qualitative advance, since current unconstrained models learn one order from data.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This survey proposes a taxonomy of Handwritten Text Recognition (HTR) methods organized by recognition granularity, dividing the field into "up to line-level" (words and lines) and "beyond line-level" (paragraphs and documents). The beyond-line category is further subdivided into attention masking, line unfolding, and unconstrained approaches. The paper provides definitions of reading order complexity, a historical narrative from handcrafted to deep-learning systems, a review of datasets and evaluation metrics, and a comparative performance analysis over the IAM benchmark. The central empirical claim, stated in Section IV-C, is that beyond-line IAM error rates are tightly clustered and that the key factor is adapting the document into a line-level structure using traditional methods.

Significance. If the taxonomy and the performance comparison were sound, this would be a useful reference for the HTR community: the reading-order-based definitions of line/paragraph/document levels are clear and the historical narrative is well grounded in the cited literature. The survey also usefully highlights the proliferation of synthetic pre-training and the lack of standardized evaluation protocols. However, the empirical conclusion about beyond-line methods rests on a comparison that is currently not reliable, because the tables mix heterogeneous evaluation protocols and contain internal inconsistencies. The taxonomic contribution is defensible, but the benchmarking analysis needs repair before the survey can serve as a trustworthy summary of state-of-the-art results.

major comments (3)
  1. [Section IV-C, Table III] The prose states that "the V AN and Origaminet, exhibit error rates of 4.7 and 4.6% respectively," but Table III lists Origaminet at 4.7% and VAN at 4.6%; the assignment is inverted, directly affecting the sentence that identifies the current state-of-the-art models.
  2. [Section IV-C, Table III] DAN is described as the open-source state-of-the-art unconstrained model, yet its CER is left as "—" in Table III, and the discussion states that DAN excludes IAM; this removes the strongest unconstrained comparison point and weakens the claim that "performance remains tightly clustered across methods," since the DAN data point is absent from the displayed cluster.
  3. [Section IV-C and Tables II/III] The central empirical conclusion that beyond-line IAM error rates are tightly clustered (4.6–6.4%) and that line-level adaptation is the key factor is drawn from Table III rows that compile CER values from papers using different IAM evaluation protocols (line-level vs. paragraph/full-page), different character sets and tokenizers, different language models/lexicons, and very different synthetic-pretraining budgets; neither Table II nor Table III includes a protocol column, and Section V-A itself concedes that heterogeneous synthetic data and tokenization practices make comparisons unfair. The cluster claim is therefore underdetermined by the evidence as presented; at minimum, the tables need explicit protocol columns and the prose must qualify the comparison.
minor comments (5)
  1. [Figure 9 caption] The caption reads "CT)" where it should read "CTC"; this typo should be corrected.
  2. [Section IV-A2] The word "sinthetic" is misspelled and should be "synthetic."
  3. [Acknowledgments] "second autor" should be "second author."
  4. [Table II] The row for C-BGRU-GRU-Att. [150] has a blank CER; either provide the value from the source or explain why it is omitted.
  5. [Throughout] The notation for the Vertical Attention Network alternates between "V AN" and "VAN"; please use one consistent form.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the survey's taxonomy and benchmark comparisons rest on external literature; the two overlapping-author citations are auxiliary and non-load-bearing.

full rationale

This is a survey with no derivation chain, so there is no equation-level circularity. The central taxonomy (Sections II-B and III) organizes methods by reading-order complexity, which is a classification choice rather than a prediction derived from the paper's own formulas. The comparative claims in Section IV-C rest on CER values taken from the original papers, which are external evidence; any concerns about heterogeneous evaluation protocols, tokenizers, synthetic pretraining budgets, or the swapped VAN/Origaminet values in Table III versus Section IV-C are correctness and reproducibility issues, not circularity. The two cited works sharing authors with this survey ([180] and [194]) are used only for an auxiliary metric discussion and a future-directions call; neither supports the taxonomy or the benchmark conclusions. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors' own prior work, and no ansatz is smuggled in via self-citation. The survey is therefore self-contained against external benchmarks for its stated organizational purpose.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The survey contributes no fitted parameters or new entities. It rests on standard probabilistic formulations plus its own organizational assumption about reading order, and on the comparability of externally reported benchmark numbers.

assumptions (4)
  • standard math The HTR task can be formulated as maximum a posteriori sequence inference: y* = argmax P(y|x) (Eq. 1), with likelihood and language-model prior (Eq. 3-4).
    Stated in Section II-C; Bayes rule and the chain rule are standard.
  • domain assumption Reading order is the defining complexity axis, so word/line, paragraph, and document levels are separated by the number of reading-order directions.
    This organizational premise is introduced in Section II-B and drives the entire taxonomy; the survey does not prove that RO is the best or only meaningful decomposition.
  • domain assumption Reported CER/WER numbers from different original papers are comparable on the IAM benchmark even though experimental setups differ.
    Used implicitly in Tables II and III and in Section IV-C's discussion; the survey itself notes tokenization and preprocessing differences but does not correct for them.
  • domain assumption The surveyed methods were evaluated under the Latin-script constraint; conclusions may not transfer to non-Latin scripts.
    Stated in Section I-A; limits scope.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Handwritten Text Recognition: A Survey." pith.science (2026). https://pith.science/paper/SWB3UVDL

@misc{pith2026250208417,
  author       = {Pith},
  title        = {Pith review of: Handwritten Text Recognition: A Survey},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SWB3UVDL}},
  note         = {Machine review of arXiv:2502.08417}
}
read the original abstract

Handwritten Text Recognition (HTR) has become an essential field within pattern recognition and machine learning, with applications spanning historical document preservation to modern data entry and accessibility solutions. The complexity of HTR lies in the high variability of handwriting, which makes it challenging to develop robust recognition systems. This survey examines the evolution of HTR models, tracing their progression from early heuristic-based approaches to contemporary state-of-the-art neural models, which leverage deep learning techniques. The scope of the field has also expanded, with models initially capable of recognizing only word-level content progressing to recent end-to-end document-level approaches. Our paper categorizes existing work into two primary levels of recognition: (1) \emph{up to line-level}, encompassing word and line recognition, and (2) \emph{beyond line-level}, addressing paragraph- and document-level challenges. We provide a unified framework that examines research methodologies, recent advances in benchmarking, key datasets in the field, and a discussion of the results reported in the literature. Finally, we identify pressing research challenges and outline promising future directions, aiming to equip researchers and practitioners with a roadmap for advancing the field.

Figures

Figures reproduced from arXiv: 2502.08417 by the authors.

Figure 1
Figure 1. Overview of the different levels of granularity in Handwritten [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Timeline of milestones in Handwritten Text Recognition (HTR). We categorize them into four levels: datasets and competitions (green), general [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Reading Order (RO) with a single direction at the line level. [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Reading Order (RO) with two directions at the paragraph level. [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]
Figure 5
Figure 5. Figure 5: Reading Order (RO) with three directions at the document level. [PITH_FULL_IMAGE:figures/full_fig_p004_5.png]
Figure 6
Figure 6. Figure 6: Overview of related fields in Handwritten Text Recognition (HTR). The figure illustrates key areas closely related to HTR, including: (1) Text [PITH_FULL_IMAGE:figures/full_fig_p005_6.png]
Figure 7
Figure 7. Figure 7: Overview of Handwritten Text Recognition (HTR) methodologies. We divide the approaches into two main groups: up-to-line and beyond line-level. [PITH_FULL_IMAGE:figures/full_fig_p006_7.png]
Figure 8
Figure 8. Figure 8: Typical pipeline before the Deep Learning era. Image from [103]. [PITH_FULL_IMAGE:figures/full_fig_p006_8.png]
Figure 9
Figure 9. Figure 9: Taxonomy of HTR approaches up to the line level. Methods are categorized into handcrafted and end-to-end approaches. Handcrafted methods rely on [PITH_FULL_IMAGE:figures/full_fig_p008_9.png]
Figure 10
Figure 10. Figure 10: Visualization of the Connectionist Temporal Classification (CTC) [PITH_FULL_IMAGE:figures/full_fig_p008_10.png]
Figure 11
Figure 11. Figure 11: Taxonomy of HTR approaches beyond line level. Methods are categorized into masking, unfolding, and unconstrained. [PITH_FULL_IMAGE:figures/full_fig_p011_11.png]
Figure 12
Figure 12. Figure 12: Illustration of beam search with beam width [PITH_FULL_IMAGE:figures/full_fig_p012_12.png]
Figure 13
Figure 13. Figure 13: Examples of handwritten text samples from various datasets commonly used in HTR research. The image showcases a range of challenges, including [PITH_FULL_IMAGE:figures/full_fig_p013_13.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4

    cs.CV 2026-07 reject novelty 4.0 of 10

    An LLM-driven closed-loop neural architecture search is applied to Arabic, Persian, and English handwriting, claiming mean test accuracies above 93% and 41–44 ms inference.

Reference graph

Works this paper leans on

206 extracted references · 70 canonical work pages · cited by 1 Pith paper

  1. [1]

    Transform- ing scholarship in the archives through handwritten text recognition: Transkribus as a case study,

    G. Muehlberger, L. Seaward, M. Terras, S. A. Oliveira, V . Bosch, M. Bryan, S. Colutto, H. D ´ejean, M. Diem, S. Fiel et al., “Transform- ing scholarship in the archives through handwritten text recognition: Transkribus as a case study,” Journal of documentation, vol. 75, no. 5, pp. 954–976, 2019

  2. [2]

    A study of children emotion and their performance while handwriting arabic char- acters using a haptic device,

    J. Zakraoui, M. Saleh, S. Al-Maadeed, and J. M. AlJa’am, “A study of children emotion and their performance while handwriting arabic char- acters using a haptic device,” Education and Information Technologies, vol. 28, no. 2, pp. 1783–1808, 2023. 17

  3. [3]

    A lexicon driven approach to handwritten word recognition for real-time applications,

    G. Kim and V . Govindaraju, “A lexicon driven approach to handwritten word recognition for real-time applications,” IEEE Trans. Pattern Anal. Mach. Intell., 1997

  4. [4]

    Off-line handwritten word recog- nition using hmm with adaptive length viterbi algorithm,

    Y . He, M.-Y . Chen, and A. Kundu, “Off-line handwritten word recog- nition using hmm with adaptive length viterbi algorithm,” ICPR, 1994

  5. [5]

    Hmm word recognition engine,

    D. Guillevic and C. Y . Suen, “Hmm word recognition engine,” 1997

  6. [6]

    Handwritten word recognition using statistics,

    T. Caesar, J. Gloger, A. Kaltenmeier, and E. Mandler, “Handwritten word recognition using statistics,” 1994

  7. [7]

    Off-line recog- nition of cursive script produced by a cooperative writer,

    H. Bunke, M. Roth, and E. G. Schukat-Talamazzini, “Off-line recog- nition of cursive script produced by a cooperative writer,” 1994

  8. [8]

    Character segmentation in handwritten words — an overview,

    M. Y . Lu, “Character segmentation in handwritten words — an overview,” Pattern Recognition, vol. 20, no. 1, pp. 77–96, 1996

Show all 206 references
  1. [9]

    A survey of methods and strategies in character segmentation,

    R. Casey and E. Lecolinet, “A survey of methods and strategies in character segmentation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 18, no. 7, pp. 690–706, 1996

  2. [10]

    An hmm-based approach for off-line unconstrained handwritten word modeling and recognition,

    M. El-Yacoubi, M. Gilloux, R. Sabourin, and C. Suen, “An hmm-based approach for off-line unconstrained handwritten word modeling and recognition,” IEEE Trans. Pattern Anal. Mach. Intell. , 1999

  3. [11]

    An introduction to hidden markov models,

    L. Rabiner and B. Juang, “An introduction to hidden markov models,” IEEE ASSP Magazine , 1986

  4. [12]

    A tutorial on hidden markov models and selected appli- cations in speech recognition,

    L. Rabiner, “A tutorial on hidden markov models and selected appli- cations in speech recognition,” Proceedings of the IEEE , 1989

  5. [13]

    Off-line handwritten word recognition using a hidden markov model type stochastic network,

    M.-Y . Chen, A. Kundu, and J. Zhou, “Off-line handwritten word recognition using a hidden markov model type stochastic network,” IEEE Trans. Pattern Anal. Mach. Intell. , 1994

  6. [14]

    Recognition of handwritten word: first and second order hidden markov model based approach,

    A. Kundu, Y . He, and P. Bahl, “Recognition of handwritten word: first and second order hidden markov model based approach,” Pattern Recognition, 1989

  7. [15]

    Deep learning,

    Y . LeCun, Y . Bengio, and G. Hinton, “Deep learning,”nature, vol. 521, no. 7553, p. 436, 2015

  8. [16]

    Gradient-based learning applied to document recognition,

    Y . Lecun, L. Bottou, Y . Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998

  9. [17]

    Long short-term memory,

    S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Comput., vol. 9, no. 8, pp. 1735–1780, 1997. [Online]. Available: https://doi.org/10.1162/neco.1997.9.8.1735

  10. [18]

    Attention is all you need,

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. u. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems , I. Guyon, U. V . Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garn...

  11. [19]

    A novel connectionist system for unconstrained handwriting recognition,

    A. Graves, M. Liwicki, S. Fern ´andez, R. Bertolami, H. Bunke, and J. Schmidhuber, “A novel connectionist system for unconstrained handwriting recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 31, no. 5, pp. 855–868, 2009

  12. [20]

    Are multidimensional recurrent layers really necessary for handwritten text recognition?

    J. Puigcerver, “Are multidimensional recurrent layers really necessary for handwritten text recognition?” in ICDAR. IEEE, 2017, pp. 67–72

  13. [21]

    Pay atten- tion to what you read: Non-recurrent handwritten text-line recognition,

    L. Kang, P. Riba, M. Rusi ˜nol, A. Forn ´es, and M. Villegas, “Pay atten- tion to what you read: Non-recurrent handwritten text-line recognition,” Pattern Recognition, vol. 129, p. 108766, 2022

  14. [22]

    Deep Neural Networks for Large V ocabulary Handwritten Text Recognition,

    T. Bluche, “Deep Neural Networks for Large V ocabulary Handwritten Text Recognition,” 2015

  15. [23]

    Transformer- based approach for joint handwriting and named entity recognition in historical document,

    A. C. Rouhou, M. Dhiaf, Y . Kessentini, and S. B. Salem, “Transformer- based approach for joint handwriting and named entity recognition in historical document,” Pattern Recognition Letters , vol. 155, pp. 128– 134, 2022

  16. [24]

    End-to-end handwritten paragraph text recognition using a vertical attention network,

    D. Coquenet, C. Chatelain, and T. Paquet, “End-to-end handwritten paragraph text recognition using a vertical attention network,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 45, no. 1, pp. 508–524, 2023

  17. [25]

    Joint line segmentation and transcription for end-to-end handwritten paragraph recognition,

    T. Bluche, “Joint line segmentation and transcription for end-to-end handwritten paragraph recognition,” in Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, December 5-10, 2016, Barcelona, Spain , 2016, pp. 838–846

  18. [26]

    A survey of text detection and recognition algorithms based on deep learning technology,

    X.-F. Wang, Z.-H. He, K. Wang, Y .-F. Wang, L. Zou, and Z.-Z. Wu, “A survey of text detection and recognition algorithms based on deep learning technology,” Neurocomputing, vol. 556, p. 126702, 2023

  19. [27]

    Advancements and challenges in handwritten text recognition: A comprehensive survey,

    W. AlKendi, F. Gechter, L. Heyberger, and C. Guyeux, “Advancements and challenges in handwritten text recognition: A comprehensive survey,” Journal of Imaging , vol. 10, no. 1, p. 18, 2024

  20. [28]

    Text recognition in the wild: A survey,

    X. Chen, L. Jin, Y . Zhu, C. Luo, and T. Wang, “Text recognition in the wild: A survey,” ACM Computing Surveys (CSUR) , vol. 54, no. 2, pp. 1–35, 2021

  21. [29]

    Automatic speech recognition and speech variability: A review,

    M. Benzeghiba, R. Mori, O. Deroo, S. Dupont, T. Erbes, D. Jouvet, L. Fissore, P. Laface, A. Mertins, C. Ris, R. Rose, V . Tyagi, and C. Wellekens, “Automatic speech recognition and speech variability: A review,” Speech Communication, 2007

  22. [30]

    Convolutional neural networks for speech recognition,

    O. Abdel-Hamid, A.-r. Mohamed, H. Jiang, L. Deng, G. Penn, and D. Yu, “Convolutional neural networks for speech recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing , 2014

  23. [31]

    Statistical methods for speech recognition,

    F. Jelinek, “Statistical methods for speech recognition,” 1997

  24. [32]

    A neural network-hidden markov model hybrid for cursive word recognition,

    S. Knerr and E. Augustin, “A neural network-hidden markov model hybrid for cursive word recognition,” 14th ICPR (Cat. No.98EX170) , 1998

  25. [33]

    Hierarchical hybrid mlp/hmm or rather mlp features for a discriminatively trained gaussian hmm: A comparison for offline handwriting recognition,

    P. Dreuw, P. Doetsch, C. Plahl, and H. Ney, “Hierarchical hybrid mlp/hmm or rather mlp features for a discriminatively trained gaussian hmm: A comparison for offline handwriting recognition,” 2011 18th IEEE International Conference on Image Processing , 2011

  26. [34]

    Off-line cursive handwriting recognition compared with on-line recognition,

    R. Seiler, M. Schenkel, and F. Eggimann, “Off-line cursive handwriting recognition compared with on-line recognition,” ICPR, 1996

  27. [35]

    Offline handwritten word recognition using a hybrid neural network and hidden markov model,

    Y . H. Tay, P.-M. Lallican, M. Khalid, C. Viard-Gaudin, and S. Knerr, “Offline handwritten word recognition using a hybrid neural network and hidden markov model,” 2001

  28. [36]

    Con- nectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,

    A. Graves, S. Fern ´andez, F. J. Gomez, and J. Schmidhuber, “Con- nectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” ICML, 2006

  29. [37]

    Empirical evaluation of gated recurrent neural networks on sequence modeling,

    J. Chung, C. Gulcehre, K. Cho, and Y . Bengio, “Empirical evaluation of gated recurrent neural networks on sequence modeling,” arXiv: Neural and Evolutionary Computing , 2014

  30. [38]

    Recurrent convolutional neural networks for text classification,

    S. Lai, L. Xu, K. Liu, and J. Zhao, “Recurrent convolutional neural networks for text classification,” AAAI, 2015

  31. [39]

    Recurrent neural networks,

    B. Hidasi, A. Karatzoglou, L. Baltrunas, and D. Tikk, “Recurrent neural networks,” Handbook on Neural Information Processing , 2020

  32. [40]

    Boosting offline handwritten text recognition in historical documents with few labeled lines,

    J. C. Aradillas, J. J. Murillo-Fuentes, and P. M. Olmos, “Boosting offline handwritten text recognition in historical documents with few labeled lines,” IEEE Access, 2021

  33. [41]

    Icdar2017 competition on handwritten text recognition on the read dataset,

    J.-A. S ´anchez, V . Romero, A. Toselli, M. Villegas, and E. Vidal, “Icdar2017 competition on handwritten text recognition on the read dataset,” 2017 14th IAPR ICDAR , 2017

  34. [42]

    Icfhr2014 competition (htrts),

    J. A. S ´anchez, V . Romero, A. Toselli, and E. Vidal, “Icfhr2014 competition (htrts),” 2014 14th ICFHR , 2014

  35. [43]

    Icfhr 2010 - competi- tions overview,

    H. E. Abed, V . M ¨argner, and M. Blumenstein, “Icfhr 2010 - competi- tions overview,” 2010

  36. [44]

    Neural machine translation by jointly learning to align and translate,

    D. Bahdanau, K. Cho, and Y . Bengio, “Neural machine translation by jointly learning to align and translate,” 2015

  37. [45]

    Attentionhtr: Handwritten text recognition based on attention encoder-decoder networks,

    D. Kass and E. Vats, “Attentionhtr: Handwritten text recognition based on attention encoder-decoder networks,” 2022

  38. [46]

    Lexicon and attention based handwritten text recognition system,

    L. Kumari, S. Singh, V . V . S. Rathore, A. Sharma, L. Kumari, S. Singh, V . V . S. Rathore, and A. Sharma, “Lexicon and attention based handwritten text recognition system,” 2022

  39. [47]

    Attention-based fully gated cnn-bgru for russian handwritten text

    A. Abdallah, M. A. Hamada, and D. Nurseitov, “Attention-based fully gated cnn-bgru for russian handwritten text.” Journal of Imaging, 2020

  40. [48]

    End-to-end handwritten paragraph text recognition using a vertical attention network,

    D. Coquenet, C. Chatelain, and T. Paquet, “End-to-end handwritten paragraph text recognition using a vertical attention network,” 2022

  41. [49]

    Evaluating sequence-to-sequence models for handwritten text recognition,

    J. Michael, R. Labahn, T. Gr ¨uning, and J. Z ¨ollner, “Evaluating sequence-to-sequence models for handwritten text recognition,” IC- DAR, 2019

  42. [50]

    An image is worth 16x16 words: Trans- formers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Trans- formers for image recognition at scale,” ArXiv, vol. abs/2010.11929, 2020

  43. [51]

    Training transformer architectures on few annotated data: an application to historical handwritten text recognition,

    K. Barrere, Y . Soullard, A. Lemaitre, and B. Co ¨uasnon, “Training transformer architectures on few annotated data: an application to historical handwritten text recognition,” IJDAR, 2024

  44. [52]

    Weakly supervised information extraction from inscrutable handwrit- ten document images,

    S. Paul, G. Madan, A. Mishra, N. Hegde, P. Kumar, and G. Aggarwal, “Weakly supervised information extraction from inscrutable handwrit- ten document images,” arXiv, 2023

  45. [53]

    Improving crnn with efficientnet-like feature extractor and multi-head attention for text recognition,

    D. V . Sang and L. T. B. Cuong, “Improving crnn with efficientnet-like feature extractor and multi-head attention for text recognition,” SoICT 2019, 2019

  46. [54]

    Dtrocr: Decoder-only transformer for optical character recognition,

    M. Fujitake, “Dtrocr: Decoder-only transformer for optical character recognition,” arXiv.org, 2023

  47. [55]

    Rethinking text line recognition models,

    D. H. Diaz, R. Ingle, S. Qin, A. Bissacco, and Y . Fujii, “Rethinking text line recognition models,” arXiv, 2021

  48. [56]

    Character-based handwritten text transcription with attention networks,

    J. Poulos and R. Valle, “Character-based handwritten text transcription with attention networks,” Neural Computing and Applications , 2021

  49. [57]

    Trocr: Transformer-based optical character recognition with pre-trained models,

    M. Li, T. Lv, J. Chen, L. Cui, Y . Lu, D. Florencio, C. Zhang, Z. Li, and F. Wei, “Trocr: Transformer-based optical character recognition with pre-trained models,” AAAI, 2023. 18

  50. [58]

    A transformer-based approach for arabic offline handwritten text recognition,

    S. Momeni and B. BabaAli, “A transformer-based approach for arabic offline handwritten text recognition,” arXiv.org, 2023

  51. [59]

    Ocformer: A transformer-based model for arabic handwritten text recognition,

    A. Mostafa, O. Mohamed, A. Ashraf, A. Elbehery, S. Jamal, G. Khoriba, and A. Ghoneim, “Ocformer: A transformer-based model for arabic handwritten text recognition,” 2021 MIUCC, 2021

  52. [60]

    Transformer for handwritten text recognition using bidirectionalfo post-decoding,

    C. Wick, J. Z ¨ollner, and T. Gr¨uning, “Transformer for handwritten text recognition using bidirectionalfo post-decoding,” ICDAR, 2021

  53. [61]

    Rescoring sequence-to-sequence models for text line recognition with ctc-prefixes,

    C. Wick, J. Z ¨ollner, and T. Gr ¨uning, “Rescoring sequence-to-sequence models for text line recognition with ctc-prefixes,” in Document Analysis Systems: 15th IAPR International Workshop, DAS 2022, La Rochelle, France, May 22–25, 2022, Proceedings . Berlin, Heidelberg: Sprin...

  54. [62]

    Dan: a segmentation-free document attention network for handwritten document recognition,

    D. Coquenet, C. Chatelain, and T. Paquet, “Dan: a segmentation-free document attention network for handwritten document recognition,” 2023

  55. [63]

    A light transformer-based architecture for handwritten text recognition,

    K. Barrere, Y . Soullard, A. Lemaitre, and B. Co ¨uasnon, “A light transformer-based architecture for handwritten text recognition,” 2022

  56. [64]

    Vilbert: Pretraining task- agnostic visiolinguistic representations for vision-and-language tasks,

    J. Lu, D. Batra, D. Parikh, and S. Lee, “Vilbert: Pretraining task- agnostic visiolinguistic representations for vision-and-language tasks,” Neural Information Processing Systems , 2019

  57. [65]

    Xlnet: Generalized autoregressive pretraining for language understanding,

    Z. Yang, Z. Dai, Y . Yang, J. G. Carbonell, R. Salakhutdinov, and Q. V . Le, “Xlnet: Generalized autoregressive pretraining for language understanding,” arXiv: Computation and Language , 2019

  58. [66]

    Sim- pler is better: Few-shot semantic segmentation with classifier weight transformer,

    Z. Lu, S. He, X. Zhu, L. Zhang, Y .-Z. Song, and T. Xiang, “Sim- pler is better: Few-shot semantic segmentation with classifier weight transformer,” 2021

  59. [67]

    Scan, attend and read: End-to-end handwritten paragraph recognition with MDLSTM attention,

    T. Bluche, J. Louradour, and R. O. Messina, “Scan, attend and read: End-to-end handwritten paragraph recognition with MDLSTM attention,” in ICDAR. IEEE, 2017, pp. 1050–1055

  60. [68]

    Dan: a segmentation-free document attention network for handwritten document recognition,

    D. Coquenet, C. Chatelain, and T. Paquet, “Dan: a segmentation-free document attention network for handwritten document recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2023

  61. [69]

    Origaminet: Weakly-supervised, segmentation-free, one-step, full page textrecognition by learning to unfold,

    M. Yousef and T. E. Bishop, “Origaminet: Weakly-supervised, segmentation-free, one-step, full page textrecognition by learning to unfold,” in CVPR, June 2020

  62. [70]

    Span: A simple predict & align network for handwritten paragraph recognition,

    D. Coquenet, C. Chatelain, and T. Paquet, “Span: A simple predict & align network for handwritten paragraph recognition,” in ICDAR, ser. Lecture Notes in Computer Science, vol. 12823, 2021, pp. 70–84

  63. [71]

    Full page handwriting recognition via image to sequence extraction,

    S. S. Singh and S. Karayev, “Full page handwriting recognition via image to sequence extraction,” in ICDAR, ser. Lecture Notes in Computer Science, J. Llad ´os, D. Lopresti, and S. Uchida, Eds., vol. 12823. Springer, 2021, pp. 55–69

  64. [72]

    Msdoctr-lite: A lite transformer for full page multi-script handwriting recognition,

    “Msdoctr-lite: A lite transformer for full page multi-script handwriting recognition,” Pattern Recognition Letters, vol. 169, pp. 28–34, 2023

  65. [73]

    The most probable string: an algorithmic study,

    C. De la Higuera and J. Oncina, “The most probable string: an algorithmic study,” Journal of Logic and Computation , vol. 24, no. 2, pp. 311–330, 2014

  66. [74]

    Advances in online handwritten recognition in the last decades,

    T. Ghosh, S. Sen, S. Obaidullah, K. Santosh, K. Roy, and U. Pal, “Advances in online handwritten recognition in the last decades,” Computer Science Review , 2022

  67. [75]

    A scalable handwritten text recognition system,

    R. R. Ingle, Y . Fujii, T. Deselaers, J. Baccash, and A. C. Popat, “A scalable handwritten text recognition system,” 2019

  68. [76]

    Online handwritten script recognition,

    A. Namboodiri and A. K. Jain, “Online handwritten script recognition,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2004

  69. [77]

    An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition,

    B. Shi, X. Bai, and C. Yao, “An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition,” CoRR, vol. abs/1507.05717, 2015. [Online]. Available: http://arxiv.org/abs/1507.05717

  70. [78]

    Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition,

    S. Fang, H. Xie, Y . Wang, Z. Mao, and Y . Zhang, “Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition,” Computer Vision and Pattern Recognition , 2021

  71. [79]

    Text recognition in the wild: A survey,

    X. Chen, L. Jin, Y . Zhu, C. Luo, and T. Wang, “Text recognition in the wild: A survey,” ACM Computing Surveys , 2021

  72. [80]

    Scene text detection and recognition: a survey,

    F. Naiemi, V . Ghods, and H. Khalesi, “Scene text detection and recognition: a survey,” Multimedia Tools and Applications , vol. 81, no. 14, pp. 20 255–20 290, 2022

  73. [81]

    Icfhr2016 handwritten keyword spotting competition (h-kws 2016),

    I. Pratikakis, K. Zagoris, B. Gatos, J. Puigcerver, A. Toselli, and E. Vidal, “Icfhr2016 handwritten keyword spotting competition (h-kws 2016),” 2016 15th ICFHR , 2016

  74. [82]

    Deep learning features for handwritten keyword spotting,

    B. Wicht, A. Fischer, and J. Hennebert, “Deep learning features for handwritten keyword spotting,” 2016 23rd ICPR , 2016

  75. [83]

    Hmm word graph based keyword spotting in handwritten document images,

    A. H. Toselli, E. Vidal, V . Romero, and V . Frinken, “Hmm word graph based keyword spotting in handwritten document images,” Information Sciences, 2016

  76. [84]

    A novel word spotting method based on recurrent neural networks,

    V . Frinken, A. Fischer, R. Manmatha, and H. Bunke, “A novel word spotting method based on recurrent neural networks,” IEEE Transac- tions on Pattern Analysis and Machine Intelligence , 2012

  77. [85]

    Docbank: A benchmark dataset for document layout analysis,

    M. Li, Y . Xu, L. Cui, S. Huang, F. Wei, Z. Li, and M. Zhou, “Docbank: A benchmark dataset for document layout analysis,” International Conference on Computational Linguistics , 2020

  78. [86]

    Doclaynet: A large human-annotated dataset for document-layout segmentation,

    B. Pfitzmann, C. Auer, M. Dolfi, A. Nassar, and P. Staar, “Doclaynet: A large human-annotated dataset for document-layout segmentation,” Knowledge Discovery and Data Mining , 2022

  79. [87]

    Document layout analysis,

    G. M. Binmakhashen and S. Mahmoud, “Document layout analysis,” ACM Computing Surveys , 2019

  80. [88]

    Docformer: End-to-end transformer for document understanding,

    S. Appalaraju, B. A. Jasani, B. Kota, Y . Xie, and R. Manmatha, “Docformer: End-to-end transformer for document understanding,” IEEE ICCV, 2021

  81. [89]

    Unidoc: Unified pretraining framework for document understanding,

    J. Gu, J. Kuen, V . I. Morariu, H. Zhao, R. Jain, N. Barmpalios, A. Nenkova, and T. Sun, “Unidoc: Unified pretraining framework for document understanding,” 2021

  82. [90]

    Document understanding dataset and evaluation (dude),

    J. V . Landeghem, R. P. Tito, L. Borchmann, M. Pietruszka, P. J’oziak, R. Powalski, D. Jurkiewicz, M. Coustaty, B. Ackaert, E. Valveny, M. B. Blaschko, S. Moens, and T. Stanislawek, “Document understanding dataset and evaluation (dude),” IEEE ICCV, 2023

  83. [91]

    Ocr-free document understanding transformer,

    G. Kim, T. Hong, M. Yim, J. Nam, J. Park, J. Yim, W. Hwang, S. Yun, D. Han, and S. Park, “Ocr-free document understanding transformer,” ECCV, 2021

  84. [92]

    Ganwriting: Content-conditioned generation of styled handwritten word images,

    L. Kang, P. Riba, Y . Wang, M. Rusi ˜nol, A. Forn ´es, and M. Villegas, “Ganwriting: Content-conditioned generation of styled handwritten word images,” Lecture Notes in Computer Science , 2020

  85. [93]

    Handwritten text generation from visual archetypes,

    V . Pippi, S. Cascianelli, and R. Cucchiara, “Handwritten text generation from visual archetypes,” CVPR, 2023

  86. [94]

    Vatr++: Choose your words wisely for handwritten text generation,

    B. Vanherle, V . Pippi, S. Cascianelli, N. Michiels, F. V . Reeth, and R. Cucchiara, “Vatr++: Choose your words wisely for handwritten text generation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2024

  87. [95]

    A multi-level perception approach to reading cursive script,

    S. Srihari and R. Bozinovic, “A multi-level perception approach to reading cursive script,” IJCAI, 1987

  88. [96]

    Reading cursive handwriting by alignment of letter prototypes,

    S. Edelman, T. Flash, and S. Ullman, “Reading cursive handwriting by alignment of letter prototypes,” International Journal of Computer Vision, 1991

  89. [97]

    Handwritten word recognition using hmm with adaptive length viterbi algorithm,

    Y . He, M.-Y . Chen, and A. Kundu, “Handwritten word recognition using hmm with adaptive length viterbi algorithm,” ICASSP, 1992

  90. [98]

    A multi-classifier combination strategy for the recognition of handwritten cursive words,

    B. Plessis, A. Sicsu, L. Heutte, E. Menu, E. Lecolinet, O. Debon, and J. Moreau, “A multi-classifier combination strategy for the recognition of handwritten cursive words,” ICDAR ’93, 1993

  91. [99]

    Improving offline handwritten text recognition with hybrid hmm/ann models,

    S. E. Boquera, M. J. C. Bleda, J. Gorbe-Moya, and F. Zamora-Mart´ınez, “Improving offline handwritten text recognition with hybrid hmm/ann models,” IEEE Transactions on Pattern Analysis and Machine Intelli- gence, 2011

  92. [100]

    Offline handwriting recognition with multidimensional recurrent neural networks,

    A. Graves and J. Schmidhuber, “Offline handwriting recognition with multidimensional recurrent neural networks,” in Advances in Neural Information Processing Systems, D. Koller, D. Schuurmans, Y . Bengio, and L. Bottou, Eds., vol. 21. Curran Associates, Inc., 2008

  93. [101]

    Large vocabulary off-line handwriting recognition: A survey,

    A. L. Koerich, R. Sabourin, and C. Y . Suen, “Large vocabulary off-line handwriting recognition: A survey,” Pattern Analysis and Applications, 2003

  94. [102]

    Holistic word recognition for handwritten historical documents,

    V . Lavrenko, T. Rath, and R. Manmatha, “Holistic word recognition for handwritten historical documents,” First International Workshop on Document Image Analysis for Libraries, 2004. Proceedings. , 2004

  95. [103]

    A survey on off-line cursive word recognition,

    A. Vinciarelli, “A survey on off-line cursive word recognition,” Pattern Recognition, 2002

  96. [104]

    Offline cursive script word recognition ? a survey,

    T. Steinherz, E. Rivlin, and N. Intrator, “Offline cursive script word recognition ? a survey,” 1999

  97. [105]

    Off-line cursive script word recognition,

    R. M. Bozinovic and S. N. Srihari, “Off-line cursive script word recognition,” 1995

  98. [106]

    Mathematical morphology and weighted least squares to correct handwriting baseline skew,

    M. Morita, J. Facon, Fl ´avio Bortolozzi, S. J. A. Garn ´es, and R. Sabourin, “Mathematical morphology and weighted least squares to correct handwriting baseline skew,” ICDAR ’99 (Cat. No.PR00318) , Sep. 1999. [Online]. Available: https://doi.org/10.1109/ICDAR.1999. 791816

  99. [107]

    Automatic reading of cursive scripts using a reading model and perceptual concepts,

    Myriam C ˆot´e, Eric Lecolinet, Mohamed Cheriet, and Ching Y . Suen, “Automatic reading of cursive scripts using a reading model and perceptual concepts,” International Journal on Document Analysis and Recognition , Feb. 1998. [Online]. Available: https: //doi.org/10.1007/s100...

  100. [108]

    OFF-LINE CURSIVE SCRIPT RECOGNITION BASED ON CONTINUOUS DENSITY HMM,

    Alessandro Vinciarelli and Juergen Luettin, “OFF-LINE CURSIVE SCRIPT RECOGNITION BASED ON CONTINUOUS DENSITY HMM,” Jan. 2004. 19

  101. [109]

    OFF-LINE UNCONSTRAINED HANDWRITTEN WORD RECOGNITION,

    Jinhai Cai and Zhi-Qiang Liu, “OFF-LINE UNCONSTRAINED HANDWRITTEN WORD RECOGNITION,” International Journal of Pattern Recognition and Artificial Intelligence , May 2000. [Online]. Available: https://doi.org/10.1142/s0218001400000180

  102. [110]

    An off-line cursive handwriting recognition system,

    A. Senior and A. J. Robinson, “An off-line cursive handwriting recognition system,” IEEE Trans. Pattern Anal. Mach. Intell. , 1998

  103. [111]

    Slant estimation algorithm for OCR systems,

    Ergina Kavallieratou, Nikos Fakotakis, and G. Kokkinakis, “Slant estimation algorithm for OCR systems,” Pattern Recognition , Dec. 2001. [Online]. Available: https://doi.org/10.1016/s0031-3203(00) 00153-9

  104. [112]

    Character segmentation in handwritten words — an overview,

    Y . Lu and M. Shridhar, “Character segmentation in handwritten words — an overview,” Pattern Recognition, 1996

  105. [113]

    Text line and word segmentation of handwritten documents,

    G. Louloudis, B. Gatos, I. Pratikakis, and C. Halatsis, “Text line and word segmentation of handwritten documents,” Pattern Recognition, 2009

  106. [114]

    A survey of methods and strategies in character segmentation,

    R. Casey and E. Lecolinet, “A survey of methods and strategies in character segmentation,” IEEE Trans. Pattern Anal. Mach. Intell., 1996

  107. [115]

    General word recognition using approximate segment-string matching,

    J. Favata, “General word recognition using approximate segment-string matching,” 1997

  108. [116]

    Offline recognition of handwritten cursive words,

    J. T. Favata and S. Srihari, “Offline recognition of handwritten cursive words,” Electronic imaging, 1992

  109. [117]

    Off line recognition of handwritten postal words using neural networks,

    C. Burges, J. Ben, J. Denker, Y . LeCun, and C. Nohl, “Off line recognition of handwritten postal words using neural networks,” Int. J. Pattern Recognit. Artif. Intell. , 1993

  110. [118]

    Handwritten word recognition using continuous density variable duration hidden markov model,

    M.-Y . Chen, A. Kundu, and S. Srihari, “Handwritten word recognition using continuous density variable duration hidden markov model,” ICASSP, 1993

  111. [119]

    A complement to variable duration hidden markov model in handwritten word recognition,

    M.-Y . Chen and A. Kundu, “A complement to variable duration hidden markov model in handwritten word recognition,” Proceedings of 1st International Conference on Image Processing , 1994

  112. [120]

    An oline cursive script recognition system using recurrent error propagation networks,

    A. Senior and F. Fallside, “An oline cursive script recognition system using recurrent error propagation networks,” 2012

  113. [121]

    Writer adaptation for handwritten word recognition using hidden markov models,

    M. Gilloux, “Writer adaptation for handwritten word recognition using hidden markov models,” 1994

  114. [122]

    Strategies for handwritten words recognition using hidden markov models,

    M. Gilloux, M. Leroux, and J. Bertille, “Strategies for handwritten words recognition using hidden markov models,” ICDAR ’93, 1993

  115. [123]

    Variable duration hidden markov model and morphological segmentation for handwritten word recog- nition,

    M.-Y . Cheii and N. Command, “Variable duration hidden markov model and morphological segmentation for handwritten word recog- nition,” 1993

  116. [124]

    Off-line handwritten word recog- nition (hwr) using a single contextual hidden markov model,

    M. Chen, A. Kundu, and J. Zhou, “Off-line handwritten word recog- nition (hwr) using a single contextual hidden markov model,” 1992

  117. [125]

    Off-line handwritten word recognition using a mixed hmm-mrf approach,

    G. Saon and A. Bela ¨ıd, “Off-line handwritten word recognition using a mixed hmm-mrf approach,” 1997

  118. [126]

    Off-line cursive handwriting recognition using hidden markov models,

    H. Bunke, M. Roth, and E. G. Schukat-Talamazzini, “Off-line cursive handwriting recognition using hidden markov models,” Pattern Recog- nition, 1995

  119. [127]

    Holistic lexicon reduction for handwritten word recognition,

    S. Madhvanath and V . Govindaraju, “Holistic lexicon reduction for handwritten word recognition,” Electronic Imaging, 1996

  120. [128]

    Global word shape processing in off-line recognition of handwriting,

    C. Parisse, “Global word shape processing in off-line recognition of handwriting,” IEEE Trans. Pattern Anal. Mach. Intell. , 1996

  121. [129]

    Pruning large lexicons using generalized word shape descriptors,

    S. Madhvanath and V . Krpasundar, “Pruning large lexicons using generalized word shape descriptors,” ICDAR, 1997

  122. [130]

    Modeling and recognition of cursive words with hidden markov models,

    W. Cho, S.-W. Lee, and J. H. Kim, “Modeling and recognition of cursive words with hidden markov models,” Pattern Recognition, 1995

  123. [131]

    Handwritten word recognition using segmentation-free hidden markov modeling and segmentation-based dynamic programming techniques,

    M. Mohamed and P. Gader, “Handwritten word recognition using segmentation-free hidden markov modeling and segmentation-based dynamic programming techniques,” IEEE Trans. Pattern Anal. Mach. Intell., 1996

  124. [132]

    Hidden markov models in handwriting recognition,

    M. Gilloux, “Hidden markov models in handwriting recognition,” 1994

  125. [133]

    Automatic reading of the literal amount of bank checks,

    T. Paquet and Y . Lecourtier, “Automatic reading of the literal amount of bank checks,” Journal of Machine Vision and Applications , 1993

  126. [134]

    An optimised minimal edit distance for hand-written word recognition,

    W. P. d. Waard, “An optimised minimal edit distance for hand-written word recognition,” Pattern Recognition Letters, 1995

  127. [135]

    Strategies for cursive script recognition using hidden markov models,

    M. Gilloux, M. Leroux, and J. M. Bertille, “Strategies for cursive script recognition using hidden markov models,” machine vision applications, 1995

  128. [136]

    Word-level optimization of dynamic programming-based handwritten word recognition algorithms,

    P. Gader and W.-T. Chen, “Word-level optimization of dynamic programming-based handwritten word recognition algorithms,” Elec- tronic Imaging, 1999

  129. [137]

    Dynamic- programming-based handwritten word recognition using the choquet fuzzy integral as the match function,

    P. D. Gader, M. A. Mohamed, and J. M. Keller, “Dynamic- programming-based handwritten word recognition using the choquet fuzzy integral as the match function,” Journal of Electronic Imaging , 1996

  130. [138]

    Handwritten word recognition for real- time applications,

    G. Kim and V . Govindaraju, “Handwritten word recognition for real- time applications,” ICDAR, 1995

  131. [139]

    Machine and human recognition of segmented characters from handwritten words,

    F. Kimura, N. Kayahara, Y . Miyake, and M. Shridhar, “Machine and human recognition of segmented characters from handwritten words,” ICDAR, 1997

  132. [140]

    Recognition of handwritten phrases as applied to street name images,

    G. Kim and V . Govindaraju, “Recognition of handwritten phrases as applied to street name images,” CVPR, 1996

  133. [141]

    Error bounds for convolutional codes and an asymptotically optimum decoding algorithm,

    A. Viterbi, “Error bounds for convolutional codes and an asymptotically optimum decoding algorithm,” IEEE Transactions on Information Theory, vol. 13, no. 2, pp. 260–269, 1967

  134. [142]

    Start, follow, read: End-to-end full-page handwriting recog- nition,

    C. Wigington, C. Tensmeyer, B. Davis, W. Barrett, B. Price, and S. Cohen, “Start, follow, read: End-to-end full-page handwriting recog- nition,” in ECCV, 2018, pp. 367–383

  135. [143]

    Offline continuous handwriting recognition using sequence to sequence neural networks,

    J. Sueiras, V . Ruiz, A. Sanchez, and J. F. Velez, “Offline continuous handwriting recognition using sequence to sequence neural networks,” Neurocomputing, 2018

  136. [144]

    Improving cnn-rnn hybrid networks for handwriting recognition,

    K. Dutta, P. Krishnan, M. Mathew, and C. V . Jawahar, “Improving cnn-rnn hybrid networks for handwriting recognition,” ICFHR, 2018

  137. [145]

    Neural machine translation by jointly learning to align and translate,

    D. Bahdanau, K. Cho, and Y . Bengio, “Neural machine translation by jointly learning to align and translate,” in ICLR, Y . Bengio and Y . LeCun, Eds., 2015

  138. [146]

    Framewise phoneme classification with bidirectional lstm and other neural network architectures,

    A. Graves and J. Schmidhuber, “Framewise phoneme classification with bidirectional lstm and other neural network architectures,” Neural Networks, vol. 18, no. 5, pp. 602–610, 2005, iJCNN 2005

  139. [147]

    Dropout improves recurrent neural networks for handwriting recognition,

    V . Pham, T. Bluche, C. Kermorvant, and J. Louradour, “Dropout improves recurrent neural networks for handwriting recognition,” in ICFHR, 2014, pp. 285–290

  140. [148]

    Gated convolutional recurrent neural networks for multilingual handwriting recognition,

    T. Bluche and R. O. Messina, “Gated convolutional recurrent neural networks for multilingual handwriting recognition,” 2017 14th IAPR ICDAR, 2017

  141. [149]

    Convolve, Attend and Spell: An Attention-based Sequence-to- Sequence Model for Handwritten Word Recognition,

    L. Kang, J. I. Toledo, P. Riba, M. Villegas, A. Forn ´es, and M. Rusi ˜nol, “Convolve, Attend and Spell: An Attention-based Sequence-to- Sequence Model for Handwritten Word Recognition,” Lecture Notes in Computer Science , 2018

  142. [150]

    Candidate fusion: Integrating language modelling into a sequence- to-sequence handwritten word recognition architecture,

    L. Kang, P. Riba, M. G. Villegas, A. Forn ´es, and M. Rusi ˜nol, “Candidate fusion: Integrating language modelling into a sequence- to-sequence handwritten word recognition architecture,” 2019

  143. [151]

    Self- attention Networks for Non-recurrent Handwritten Text Recognition,

    R. d’Arce, T. Norton, S. Hannuna, and N. Cristianini, “Self- attention Networks for Non-recurrent Handwritten Text Recognition,” in Frontiers in Handwriting Recognition , U. Porwal, A. Forn ´es, and F. Shafait, Eds. Cham: Springer International Publishing, 2022, vol. 13639, pp...

  144. [152]

    An efficient end-to-end neural model for handwritten text recognition,

    A. Chowdhury and L. Vig, “An efficient end-to-end neural model for handwritten text recognition,” in British Machine Vision Conference ,

  145. [153]

    A learning algorithm for continually running fully recurrent neural networks,

    R. J. Williams and D. Zipser, “A learning algorithm for continually running fully recurrent neural networks,” Neural Computation, vol. 1, no. 2, pp. 270–280, 1989

  146. [154]

    Character-based handwritten text transcription with attention networks,

    J. Poulos and R. Valle, “Character-based handwritten text transcription with attention networks,” Neural computing & applications (Print) , 2017

  147. [155]

    Joint ctc-attention based end-to- end speech recognition using multi-task learning,

    S. Kim, T. Hori, and S. Watanabe, “Joint ctc-attention based end-to- end speech recognition using multi-task learning,” in ICASSP, 2017, pp. 4835–4839

  148. [156]

    Handwritten document recog- nition using pre-trained vision transformers,

    D. Parres, D. Anitei, and R. Paredes, “Handwritten document recog- nition using pre-trained vision transformers,” in Document Analysis and Recognition - ICDAR 2024 . Cham: Springer Nature Switzerland, 2024, pp. 173–190

  149. [157]

    The iam-database: an english sentence database for offline handwriting recognition,

    U.-V . Marti and H. Bunke, “The iam-database: an english sentence database for offline handwriting recognition,” International Journal on Document Analysis and Recognition , 2002

  150. [158]

    Handwritten mail classification experiments with the rimes database,

    C. Kermorvant and J. Louradour, “Handwritten mail classification experiments with the rimes database,” in ICFHR. IEEE Computer Society, 2010, pp. 241–246. [Online]. Available: https://doi.org/10. 1109/ICFHR.2010.45

  151. [159]

    Building a volunteer community: Results and findings from transcribe bentham,

    T. Causer and V . Wallace, “Building a volunteer community: Results and findings from transcribe bentham,” Digital Humanities Quarterly , vol. 6, no. 2, 2012

  152. [160]

    Scherrer, Verzeichniss der Handschriften der Stiftsbibliothek von St

    G. Scherrer, Verzeichniss der Handschriften der Stiftsbibliothek von St. Gallen. Halle, 1875

  153. [161]

    The RODRIGO database,

    N. Serrano, F. Castro, and A. Juan, “The RODRIGO database,” in Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC’10) , N. Calzolari, K. Choukri, B. Maegaard, J. Mariani, J. Odijk, S. Piperidis, M. Rosner, and D. Tapias, Eds. Vallett...

  154. [162]

    Icfhr2016 competition on handwritten text recognition on the read dataset,

    J. A. S ´anchez, V . Romero, A. H. Toselli, and E. Vidal, “Icfhr2016 competition on handwritten text recognition on the read dataset,” in ICFHR, 2016, pp. 630–635

  155. [163]

    Lexicon-free handwritten word spotting using character hmms,

    A. Keller, V . Frinken, and H. Bunke, “Lexicon-free handwritten word spotting using character hmms,” Pattern Recognition Letters - PRL , vol. 33, p. 934–942, 05 2012

  156. [164]

    The esposalles database: An ancient marriage license corpus for off-line handwriting recognition,

    V . Romero, A. Forn ´eS, N. Serrano, J. A. S ´aNchez, A. H. Toselli, V . Frinken, E. Vidal, and J. Llad ´oS, “The esposalles database: An ancient marriage license corpus for off-line handwriting recognition,” Pattern Recogn., vol. 46, no. 6, p. 1658–1669, Jun. 2013

  157. [165]

    Generating synthetic data for text recognition,

    P. Krishnan and C. V . Jawahar, “Generating synthetic data for text recognition,” CoRR, vol. abs/1608.04224, 2016. [Online]. Available: http://arxiv.org/abs/1608.04224

  158. [166]

    Unsuper- vised adaptation for synthetic-to-real handwritten word recognition,

    L. Kang, M. Rusinol, A. Fornes, P. Riba, and M. Villegas, “Unsuper- vised adaptation for synthetic-to-real handwritten word recognition,” workshop on applications of computer vision , 2020

  159. [167]

    Brown corpus manual,

    W. N. Francis and H. Kucera, “Brown corpus manual,” Department of Linguistics, Brown University, Providence, Rhode Island, US, Tech. Rep., 1979

  160. [168]

    Daniel: A fast document attention network for information extraction and labelling of handwritten documents,

    T. Constum, P. Tranouez, and T. Paquet, “Daniel: A fast document attention network for information extraction and labelling of handwritten documents,” 2024. [Online]. Available: https://arxiv.org/ abs/2407.09103

  161. [169]

    Pointer sentinel mixture models,

    S. Merity, C. Xiong, J. Bradbury, and R. Socher, “Pointer sentinel mixture models,” 2016

  162. [170]

    The significance of reading order in document recognition and its evaluation,

    C. Clausner, S. Pletschacher, and A. Antonacopoulos, “The significance of reading order in document recognition and its evaluation,” inICDAR. IEEE, 2013, pp. 688–692

  163. [171]

    Europeana newspapers OCR workflow evaluation,

    S. Pletschacher, C. Clausner, and A. Antonacopoulos, “Europeana newspapers OCR workflow evaluation,” in Proceedings of the 3rd In- ternational Workshop on Historical Document Imaging and Processing. New York, NY , USA: Association for Computing Machinery, 2015, p. 39–46

  164. [172]

    Flexible char- acter accuracy measure for reading-order-independent evaluation,

    C. Clausner, S. Pletschacher, and A. Antonacopoulos, “Flexible char- acter accuracy measure for reading-order-independent evaluation,” Pat- tern Recognit. Lett. , vol. 131, pp. 390–397, 2020

  165. [173]

    How much data do you need? about the creation of a ground truth for black letter and the effectiveness of neural OCR,

    P. B. Str ¨obel, S. Clematide, and M. V olk, “How much data do you need? about the creation of a ground truth for black letter and the effectiveness of neural OCR,” in Proceedings of the Twelfth Language Resources and Evaluation Conference . Marseille, France: European Languag...

  166. [174]

    ICDAR2017 competition on recognition of documents with complex layouts- RDCL2017,

    C. Clausner, A. Antonacopoulos, and S. Pletschacher, “ICDAR2017 competition on recognition of documents with complex layouts- RDCL2017,” in 2017 14th ICDAR, vol. 1. IEEE, 2017, pp. 1404–1410

  167. [175]

    ICDAR2019 competition on recognition of documents with complex layouts - RDCL2019,

    ——, “ICDAR2019 competition on recognition of documents with complex layouts - RDCL2019,” in ICDAR. IEEE, 2019, pp. 1521– 1526

  168. [176]

    Training full-page handwritten text recognition models without annotated line breaks,

    C. Tensmeyer and C. Wigington, “Training full-page handwritten text recognition models without annotated line breaks,” in ICDAR. IEEE, 2019, pp. 1–8

  169. [177]

    Towards end-to-end unified scene text detection and layout analysis,

    S. Long, S. Qin, D. Panteleev, A. Bissacco, Y . Fujii, and M. Raptis, “Towards end-to-end unified scene text detection and layout analysis,” in CVPR, 2022, pp. 1049–1059

  170. [178]

    ICDAR2017 competition on handwritten text recognition on the READ dataset,

    J. A. S ´anchez, V . Romero, A. H. Toselli, M. Villegas, and E. Vidal, “ICDAR2017 competition on handwritten text recognition on the READ dataset,” in 2017 14th ICDAR) , vol. 01, 2017, pp. 1383–1388

  171. [179]

    Bleu: a method for automatic evaluation of machine translation,

    K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu, “Bleu: a method for automatic evaluation of machine translation,” in Proceedings of the 40th annual meeting of the Association for Computational Linguistics , 2002, pp. 311–318

  172. [180]

    End- to-end page-level assessment of handwritten text recognition,

    E. Vidal, A. H. Toselli, A. R ´ıos-Vila, and J. Calvo-Zaragoza, “End- to-end page-level assessment of handwritten text recognition,” Pattern Recognition, vol. 142, p. 109695, 2023

  173. [181]

    Boosting modern and historical handwritten text recognition with deformable convolutions,

    S. Cascianelli, M. Cornia, L. Baraldi, and R. Cucchiara, “Boosting modern and historical handwritten text recognition with deformable convolutions,” IJDAR, vol. 25, pp. 207 – 217, 2022

  174. [182]

    Neural Networks for Handwrit- ing Recognition,

    M. Liwicki, A. Graves, and H. Bunke, “Neural Networks for Handwrit- ing Recognition,” Computational Intelligence Paradigms in Advanced Pattern Classification, 2012

  175. [183]

    Improvements in RWTH’s System for Off-Line Handwriting Recognition,

    M. Kozielski, P. Doetsch, and H. Ney, “Improvements in RWTH’s System for Off-Line Handwriting Recognition,” ICDAR, 2013

  176. [184]

    Open vocabulary handwriting recognition using combined word-level and character-level language models,

    M. Kozielski, D. Rybach, S. Hahn, R. Schl ¨uter, and H. Ney, “Open vocabulary handwriting recognition using combined word-level and character-level language models,” ICASSP, 2013

  177. [185]

    Fast and Robust Training of Recurrent Neural Networks for Offline Handwriting Recognition,

    P. Doetsch, M. Kozielski, and H. Ney, “Fast and Robust Training of Recurrent Neural Networks for Offline Handwriting Recognition,” ICFHR, 2014

  178. [186]

    Sequence-discriminative training of recurrent neural networks,

    P. V oigtlaender, P. Doetsch, S. Wiesler, R. Schl ¨uter, and H. Ney, “Sequence-discriminative training of recurrent neural networks,” ICASSP, 2015

  179. [187]

    Handwriting Recognition with Large Multidimensional Long Short-Term Memory Recurrent Neural Networks,

    P. V oigtlaender, P. Doetsch, and H. Ney, “Handwriting Recognition with Large Multidimensional Long Short-Term Memory Recurrent Neural Networks,” 2016 15th ICFHR , 2016

  180. [188]

    Simultaneous script identifi- cation and handwriting recognition via multi-task learning of recurrent neural networks,

    Z. Chen, Y . Wu, F. Yin, and C.-L. Liu, “Simultaneous script identifi- cation and handwriting recognition via multi-task learning of recurrent neural networks,” in 2017 14th IAPR ICDAR , vol. 01, 2017, pp. 525– 530

  181. [189]

    Boosting the Deep Mul- tidimensional Long-Short-Term Memory Network for Handwritten Recognition Systems,

    D. Castro, B. Bezerra, and M. Valenc ¸a, “Boosting the Deep Mul- tidimensional Long-Short-Term Memory Network for Handwritten Recognition Systems,” ICFHR, 2018

  182. [190]

    Adaptive Context- aware Reinforced Agent for Handwritten Text Recognition,

    L. Gui, X. Liang, X. Chang, and A. Hauptmann, “Adaptive Context- aware Reinforced Agent for Handwritten Text Recognition,” British Machine Vision Conference, 2018

  183. [191]

    Word spotting and recog- nition using deep embedding,

    P. Krishnan, K. Dutta, and C. V . Jawahar, “Word spotting and recog- nition using deep embedding,” 2018 13th DAS , 2018

  184. [192]

    End-to-end sequence labeling via convolutional recurrent neural network with a connectionist temporal classification layer,

    X. Huang, L. Qiao, W. Yu, J. Li, and Y . Ma, “End-to-end sequence labeling via convolutional recurrent neural network with a connectionist temporal classification layer,” International Journal of Computational Intelligence Systems , vol. 13, pp. 341–351, 2020. [Online]. Availa...

  185. [193]

    HTR-VT: Handwritten text recognition with vision transformer,

    Y . Li, D. Chen, T. Tang, and X. Shen, “HTR-VT: Handwritten text recognition with vision transformer,” Pattern Recognition, 2024

  186. [194]

    On the generalization of handwritten text recognition models,

    C. Garrido-Munoz and J. Calvo-Zaragoza, “On the generalization of handwritten text recognition models,” 2024. [Online]. Available: https://arxiv.org/abs/2411.17332

  187. [195]

    Fine-tuning is a surprisingly effective domain adaptation baseline in handwriting recognition,

    J. Koh ´ut and M. Hradis, “Fine-tuning is a surprisingly effective domain adaptation baseline in handwriting recognition,” Lecture Notes in Computer Science, 2023

  188. [196]

    Metahtr: Towards writer-adaptive handwritten text recognition,

    A. K. Bhunia, S. Ghose, A. Kumar, P. N. Chowdhury, A. Sain, and Y .-Z. Song, “Metahtr: Towards writer-adaptive handwritten text recognition,” Computer Vision and Pattern Recognition , 2021

  189. [197]

    Paligemma: A versatile 3b vlm for transfer,

    L. B. et al., “Paligemma: A versatile 3b vlm for transfer,” ArXiv, vol. abs/2407.07726, 2024

  190. [198]

    Visual instruction tuning,

    H. Liu, C. Li, Q. Wu, and Y . J. Lee, “Visual instruction tuning,” ArXiv, vol. abs/2304.08485, 2023

  191. [199]

    Layoutlm: Pre-training of text and layout for document image understanding,

    Y . Xu, M. Li, L. Cui, S. Huang, F. Wei, and M. Zhou, “Layoutlm: Pre-training of text and layout for document image understanding,” in Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining , ser. KDD ’20. Association for Computing Mac...

  192. [200]

    Ocr-free document understanding transformer,

    G. Kim, T. Hong, M. Yim, J. Nam, J. Park, J. Yim, W. Hwang, S. Yun, D. Han, and S. Park, “Ocr-free document understanding transformer,” in ECCV, 2021

  193. [201]

    Self-supervised learning: The dark matter of intelligence

    Y . LeCun, “Self-supervised learning: The dark matter of intelligence.” [Online]. Available: https://ai.facebook.com/blog/ self-supervised-learning-the-dark-matter-of-intelligence/

  194. [202]

    Know your self-supervised learn- ing: A survey on image-based generative and discriminative training,

    U. Ozbulak, H. J. Lee, B. Boga, E. T. Anzaku, H. Park, A. Van Messem, W. De Neve, and J. Vankerschaver, “Know your self-supervised learn- ing: A survey on image-based generative and discriminative training,” arXiv preprint arXiv:2305.13689 , 2023

  195. [203]

    BERT: pre-training of deep bidirectional transformers for language understanding,

    J. Devlin, M. Chang, K. Lee, and K. Toutanova, “BERT: pre-training of deep bidirectional transformers for language understanding,” in NAACL-HLT 2019, Vol. 1, J. Burstein, C. Doran, and T. Solorio, Eds. Association for Computational Linguistics, 2019, pp. 4171–4186

  196. [204]

    Beit: BERT pre-training of image transformers,

    H. Bao, L. Dong, and F. Wei, “Beit: BERT pre-training of image transformers,” in 10th ICLR, Apr 2022, Virtual, France , 2022

  197. [205]

    Sequence-to-sequence contrastive learn- ing for text recognition,

    A. Aberdam, R. Litman, S. Tsiper, O. Anschel, R. Slossberg, S. Mazor, R. Manmatha, and P. Perona, “Sequence-to-sequence contrastive learn- ing for text recognition,” in CVPR 2021. Computer Vision Foundation / IEEE, 2021, pp. 15 302–15 312

  198. [2018]

    Available: https://api.semanticscholar.org/CorpusID: 49907144

    [Online]. Available: https://api.semanticscholar.org/CorpusID: 49907144

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.