REVIEW 3 major objections 4 minor 60 references
Relational Contrastive Learning and Masked Image Modeling for Scene Text Recognition
T0 review · 3 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Self-supervised pretraining for scene text recognition can be reframed as learning relations among textual elements, and the proposed RCMSTR framework reports state-of-the-art results across frozen-feature, fine-tuning, semi-supervised…
desk verdict A solid self-supervised STR recipe with consistent gains, but the headline margin over DiG rests partly on test-set-tuned hyperparameters and the code is not out yet. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The framework has two interacting branches. The relational contrastive branch (RCL) builds on a MoCo-style momentum contrastive base, but replaces whole-image instance discrimination with relational contrastive learning that uses KL divergence to align similarity distributions. A permutation module splits each image horizontally into N patches and shuffles patches across M images to create new 'texts' such as 'justice' from 'justify' and 'notice', enriching the finite relation set. Three predictors pool encoder features into frame, subword, and word levels, and the model pulls together positive pairs at each level while imposing KL-based consistency between frame-subword and subword-word relations. The masked image modeling branch (MIM) masks both random patches (for local strokes) and horizontal blocks (for whole characters), reconstructing normalized RGB pixels with an L2 loss. A decoupling design feeds only unmasked views to the contrastive branch and masked views to the reconstruction branch, sharing encoder weights, which the paper argues avoids the distribution mismatch that makes naive CL+MIM coupling unstable. For CNN backbones, all standard convolutions are replaced with sparse convolutions during pretraining, following ConvNeXt V2, so masked inputs can be processed efficiently.
What would settle it
Pre-train RCMSTR and a baseline on a dataset of rotated, curved, or vertical text and compare frozen-feature recognition accuracy; if RCMSTR's margin shrinks or reverses on such data, the horizontal-text assumption is the cause.
Extended reading notes
Core claim
The central claim is that arranging scene text into its natural hierarchy — frames, subwords, and words — and treating the relations among those levels as self-supervised labels yields better representations than treating whole text images as contrastive instances or applying generic natural-image MIM. RCMSTR outperforms the previous state-of-the-art self-supervised STR method DiG on frozen ViT features by 3.32 average accuracy points (58.69 vs 55.37 across 12 datasets), and on CNN features it reaches 53.09 average accuracy with an attention decoder compared with 42.79 for the best RCLSTR variant without MIM. The paper also reports 80.77% average accuracy after supervised fine-tuning on ViT and 62.35% average accuracy with only 1% labeled data, both ahead of the compared baselines.
Load-bearing premise
Everything rests on the text being horizontal and left-to-right: the permutation module shuffles horizontal patches and the MIM masks horizontal spans, so for rotated, curved, vertical, or arbitrarily oriented text the rearranged images break character order and the block masks stop matching character-level context.
Editorial extensions
If this is right
- Frozen-feature accuracy improves monotonically as each module is added (enriching relations, intra-hierarchy, inter-hierarchy, then MIM), so each component carries part of the gain.
- The same pretrained weights help downstream tasks beyond recognition, including text segmentation on TextSeg, where RCMSTR-ViT-Small reaches 83.9 IoU versus 81.1 from scratch.
- With only 1% of labels, RCMSTR reaches 62.35% average accuracy versus 46.07% for a randomly initialized supervised baseline, so pretraining substantially reduces labeling cost.
- The method transfers to Chinese documents and handwritten English, where horizontal structure and multi-granularity still hold.
- Because the MIM branch works with CNN encoders via sparse convolutions, the gains are not tied to ViT only.
Reading between the lines
- If the horizontal-text assumption is the limiting factor, a rotation-aware or curvature-aware variant of the permutation and block-masking modules could plausibly extend the same relational pretraining to arbitrary-orientation text, but the paper does not test this.
- The decoupling design may generalize beyond STR: any contrastive-plus-MIM pipeline that currently feeds masked positives into the contrastive branch could adopt the unmasked-only contrastive stream and weight-shared reconstruction stream.
- Block masking along the reading direction acts as a character-level pretext task without any character annotation, suggesting a cheap way to inject language-like context into OCR pretraining on unlabeled corpora.
- The reported gains are averages over 12 datasets with mixed difficulty; the largest relative improvements appear on occluded and perspective benchmarks, hinting that relational pretraining especially helps hard cases.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents RCMSTR, a self-supervised pretraining framework for scene text recognition that combines relational contrastive learning (RCL) with masked image modeling (MIM). The RCL branch enriches textual relations by dividing images into horizontal patches and permuting them across images, then applies contrastive losses at frame, subword, and word levels plus KL-based inter-hierarchical consistency losses. The MIM branch uses random patch masking together with a horizontal block-masking strategy, and the two branches are integrated through a decoupling design in which masked images are used only for reconstruction. Sparse convolutions are introduced to make MIM compatible with CNN encoders. The method is pretrained on SynthText and evaluated on twelve scene-text benchmarks under frozen-feature, fine-tuning, and semi-supervised protocols, as well as a text-segmentation transfer task; the authors report consistent improvements over existing self-supervised STR methods.
Significance. If the reported gains are robust, the paper would be a useful contribution to self-supervised STR, showing that text-specific inductive biases such as rearrangement, hierarchical contrastive learning, and horizontal block masking improve representations for both CNN and ViT encoders. The paper contains extensive ablations, sequential module analysis in Table I, and evaluation across several downstream tasks, which are strengths. It also discloses the relation to the authors' earlier RCLSTR work. However, the headline SOTA claim is currently supported by comparisons to a small set of baselines, with MIM hyperparameters selected on the evaluation benchmarks themselves and no seed variance, so the true magnitude and statistical reliability of the gains are not yet established.
major comments (3)
- [Section IV-H, Figures 7-8, Table I] The default MIM hyperparameters (mask ratio 0.7 and one horizontal masked block) are selected from Figures 7 and 8, whose y-axis is the average representation accuracy on the first seven test datasets (IIIT5K through CUTE). These same test datasets are used in the headline ViT comparison in Table I, where RCMSTR (58.69) beats the reproduced DiG-dagger baseline (55.37) by 3.32 points. Because DiG-dagger is evaluated with fixed hyperparameters while RCMSTR's MIM hyperparameters are tuned on these test sets, and because no validation split, error bars, or code are provided, part or all of the gap could be a selection artifact. Please choose mask ratio and block number on a held-out validation split, or report a hyperparameter sensitivity analysis with standard deviations over multiple seeds, and rerun the main Table I comparisons under that protocol.
- [Section III-C2, Table VI] The decoupling design is presented as a key contribution, stated to effectively integrate RCL and MIM and to mitigate instability caused by feeding masked images into the CL online encoder, but no ablation compares the decoupled integration with the coupled alternative (the DiG-style design in which masked images are also used for CL). Without such an ablation, the claim that the decoupling design itself is responsible for the gains of RCMSTR over a naive integration is not empirically supported. Please add a coupled-variant comparison under identical mask ratio, loss weights, training iterations, and encoder architecture.
- [Section IV-A, Tables I-III] The experimental section omits several details needed to interpret the listed accuracies: the number of pre-training epochs, optimizer and learning-rate schedule, batch size, MoCo queue size K, number of GPUs, and the exact composition of the SynthText pretraining set. In addition, all reported numbers are single runs with no standard deviation over seeds. Because the main comparisons in Tables I-III include average differences as small as a few points, the paper should report the mean and variance over at least three pretraining seeds, or justify why single-run results are sufficient for the claims made.
minor comments (4)
- [Section IV-E, Tables I-IV] The paper states that the method is conditioned solely on the assumption that text is horizontal, yet Tables I-III include datasets with curved and perspective text (CUTE80, SVTP, TT, CTW). Please clarify whether these datasets are intended to test generalization beyond the horizontal assumption and how the relational modules behave when the reading order is not strictly left-to-right.
- [Section III-B2, Equation (5)] The claim that averaging projected features into T=4 subword segments approximates morphological units such as roots and affixes is an assumption; Section III-B2 states this as a default without evidence. A short analysis or reference supporting the choice T=4 would help the reader evaluate the hierarchical design.
- [Section IV-A] The data augmentation description says the authors follow SeqCLR, but the augmentation probabilities and the exact masking-augmentation order are not listed. Please provide the full augmentation configuration for reproducibility.
- [Section IV-E, Table IV] There is a stray period after 'summarized in Table IV. .' in the text; please fix the typo.
Circularity Check
No significant circularity: RCMSTR's losses are self-supervised objectives with external downstream evaluation, and the prior RCLSTR self-citation is disclosed and not load-bearing.
full rationale
RCMSTR's derivation chain is not circular. The RCL loss (Eqs. 1-7) and MIM loss (Eq. 8) are standard self-supervised objectives with pixel or feature targets, and no fitted constant enters the loss definition; the learned representation is then evaluated by separate downstream decoders (Tables I-III). The prior RCLSTR self-citation ([22]) is disclosed in Section I and used only as a component/baseline; the new MIM, decoupling, and CNN compatibility claims are supported by ablations and by a reproduced DiG† baseline (Table I: 58.69 vs 55.37) that does not depend on the citation. Selecting the mask ratio and number of masked blocks from Figures 7-8, whose y-axis averages the first seven test sets, is a test-set tuning concern rather than a constructional circularity: the reported gains also appear on the five datasets not used in that selection (e.g., CTW 48.28 vs 44.97, TT 50.39 vs 46.93, WOST 44.54 vs 42.51), and no equation or prediction reduces to a fitted quantity by definition.
Assumptions & free parameters
free parameters (9)
- alpha (KL loss weight) =
0.3
- beta (RCL loss weight) =
0.1
- T (number of subword segments) =
4
- N (horizontal patches per image) =
2
- M (images per permutation group) =
2
- Patch mask ratio =
0.7
- Number of horizontal masked blocks =
1
- Temperatures tau_info and tau_kl =
unspecified
- MoCo queue size K =
unspecified
assumptions (5)
- domain assumption Text in target images is horizontal and left-to-right
- domain assumption Pre-training on SynthText transfers to real scene text benchmarks
- domain assumption Relational similarity-distribution consistency (KL over negatives) is a meaningful self-supervised objective
- ad hoc to paper Averaging features into T=4 subword bins approximates morphological units such as roots and affixes
- domain assumption Sparse convolution equals standard convolution on unmasked inputs
Cite this review
Pith. "Pith review of Relational Contrastive Learning and Masked Image Modeling for Scene Text Recognition." pith.science (2026). https://pith.science/paper/6OQMY3QE
@misc{pith2026241111219,
author = {Pith},
title = {Pith review of: Relational Contrastive Learning and Masked Image Modeling for Scene Text Recognition},
year = {2026},
howpublished = {\url{https://pith.science/paper/6OQMY3QE}},
note = {Machine review of arXiv:2411.11219}
}
read the original abstract
Context-aware methods have achieved remarkable advancements in supervised scene text recognition by leveraging semantic priors from words. Considering the heterogeneity of text and background in STR, we propose that such contextual priors can be reinterpreted as the relations between textual elements, serving as effective self-supervised labels for representation learning. However, textual relations are restricted to the finite size of the dataset due to lexical dependencies, which causes over-fitting problem, thus compromising the representation quality. To address this, our work introduces a unified framework of Relational Contrastive Learning and Masked Image Modeling for STR (RCMSTR), which explicitly models the enriched textual relations. For the RCL branch, we first introduce the relational rearrangement module to cultivate new relations on the fly. Based on this, we further conduct relational contrastive learning to model the intra- and inter-hierarchical relations for frames, sub-words and words. On the other hand, MIM can naturally boost the context information via masking, where we find that the block masking strategy is more effective for STR. For the effective integration of RCL and MIM, we also introduce a novel decoupling design aimed at mitigating the impact of masked images on contrastive learning. Additionally, to enhance the compatibility of MIM with CNNs, we propose the adoption of sparse convolutions and directly sharing the weights with dense convolutions in training. The proposed RCMSTR demonstrates superior performance in various evaluation protocols for different STR-related downstream tasks, outperforming the existing state-of-the-art self-supervised STR techniques. Ablation studies and qualitative experimental results further validate the effectiveness of our method. The code and pre-trained models will be available at https://github.com/ThunderVVV/RCMSTR .
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Unsupervised feature learning via non-parametric instance discrimination,
Z. Wu, Y . Xiong, S. X. Yu, and D. Lin, “Unsupervised feature learning via non-parametric instance discrimination,” in Proceedings of the IEEE conference on computer vision and pattern recognition . IEEE, 2018, pp. 3733–3742
work page 2018
-
[2]
Representation learning with contrastive predictive coding,
A. v. d. Oord, Y . Li, and O. Vinyals, “Representation learning with contrastive predictive coding,” arXiv preprint arXiv:1807.03748 , 2018
arXiv 2018
-
[3]
Momentum contrast for unsupervised visual representation learning,
K. He, H. Fan, Y . Wu, S. Xie, and R. Girshick, “Momentum contrast for unsupervised visual representation learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 9729–9738
2020
-
[4]
A simple framework for contrastive learning of visual representations,
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in International conference on machine learning . PMLR, 2020, pp. 1597–1607
2020
-
[5]
Beit: Bert pre-training of image transformers,
H. Bao, L. Dong, S. Piao, and F. Wei, “Beit: Bert pre-training of image transformers,” arXiv preprint arXiv:2106.08254 , 2021
arXiv 2021
-
[6]
Masked au- toencoders are scalable vision learners,
K. He, X. Chen, S. Xie, Y . Li, P. Doll ´ar, and R. Girshick, “Masked au- toencoders are scalable vision learners,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2022, pp. 16 000–16 009
2022
-
[7]
Simmim: A simple framework for masked image modeling,
Z. Xie, Z. Zhang, Y . Cao, Y . Lin, J. Bao, Z. Yao, Q. Dai, and H. Hu, “Simmim: A simple framework for masked image modeling,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. IEEE COMPUTER SOC, 2022, pp. 9653–9663
work page 2022
-
[8]
S. Fang, H. Xie, Y . Wang, Z. Mao, and Y . Zhang, “Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 7098–7107
work page 2021
Show all 60 references
-
[11]
Sequence-to-sequence contrastive learn- ing for text recognition,
A. Aberdam, R. Litman, S. Tsiper, O. Anschel, R. Slossberg, S. Mazor, R. Manmatha, and P. Perona, “Sequence-to-sequence contrastive learn- ing for text recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 15 302–15 312
2021
-
[12]
Perceiving stroke-semantic context: Hierarchical contrastive learning for robust scene text recognition,
H. Liu, B. Wang, Z. Bao, M. Xue, S. Kang, D. Jiang, Y . Liu, and B. Ren, “Perceiving stroke-semantic context: Hierarchical contrastive learning for robust scene text recognition,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 2, 2022, pp. 1702– 1710
2022
-
[13]
Reading and writing: Discriminative and generative modeling for self-supervised text recognition,
M. Yang, M. Liao, P. Lu, J. Wang, S. Zhu, H. Luo, Q. Tian, and X. Bai, “Reading and writing: Discriminative and generative modeling for self-supervised text recognition,” in Proceedings of the 30th ACM International Conference on Multimedia , 2022, pp. 4214–4223
2022
-
[14]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929 , 2020
2010 arXiv
-
[15]
Maskocr: Text recognition with masked encoder-decoder pretraining,
P. Lyu, C. Zhang, S. Liu, M. Qiao, Y . Xu, L. Wu, K. Yao, J. Han, E. Ding, and J. Wang, “Maskocr: Text recognition with masked encoder-decoder pretraining,” arXiv preprint arXiv:2206.00311 , 2022
2022 arXiv
-
[16]
Towards accurate scene text recognition with semantic reasoning networks,
D. Yu, X. Li, C. Zhang, T. Liu, J. Han, J. Liu, and E. Ding, “Towards accurate scene text recognition with semantic reasoning networks,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 12 113–12 122
2020
-
[17]
Seed: Semantics enhanced encoder-decoder framework for scene text recognition,
Z. Qiao, Y . Zhou, D. Yang, Y . Zhou, and W. Wang, “Seed: Semantics enhanced encoder-decoder framework for scene text recognition,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 13 528–13 537
2020
-
[18]
On vocabulary reliance in scene text recognition,
Z. Wan, J. Zhang, L. Zhang, J. Luo, and C. Yao, “On vocabulary reliance in scene text recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 11 425–11 434
2020
-
[19]
Ressl: Relational self-supervised learning with weak augmentation,
M. Zheng, S. You, F. Wang, C. Qian, C. Zhang, X. Wang, and C. Xu, “Ressl: Relational self-supervised learning with weak augmentation,” Advances in Neural Information Processing Systems , vol. 34, pp. 2543– 2555, 2021
2021
-
[20]
Synthetic data for text localisation in natural images,
A. Gupta, A. Vedaldi, and A. Zisserman, “Synthetic data for text localisation in natural images,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 2315–2324
2016
-
[21]
Rethink- ing text segmentation: A novel dataset and a text-specific refinement approach,
X. Xu, Z. Zhang, Z. Wang, B. Price, Z. Wang, and H. Shi, “Rethink- ing text segmentation: A novel dataset and a text-specific refinement approach,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 12 045–12 055
2021
-
[22]
Relational contrastive learning for scene text recognition,
J. Zhang, T. Lin, Y . Xu, K. Chen, and R. Zhang, “Relational contrastive learning for scene text recognition,” in Proceedings of the 31st ACM International Conference on Multimedia , 2023, pp. 5764–5775
2023
-
[23]
Unsupervised learning of visual features by contrasting cluster assign- ments,
M. Caron, I. Misra, J. Mairal, P. Goyal, P. Bojanowski, and A. Joulin, “Unsupervised learning of visual features by contrasting cluster assign- ments,” Advances in neural information processing systems , vol. 33, pp. 9912–9924, 2020
2020
-
[24]
Bootstrap your own latent-a new approach to self-supervised learning,
J.-B. Grill, F. Strub, F. Altch ´e, C. Tallec, P. Richemond, E. Buchatskaya, C. Doersch, B. Avila Pires, Z. Guo, M. Gheshlaghi Azar et al. , “Bootstrap your own latent-a new approach to self-supervised learning,” Advances in neural information processing systems, vol. 33, pp. ...
2020
-
[25]
Exploring simple siamese representation learning,
X. Chen and K. He, “Exploring simple siamese representation learning,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 15 750–15 758
2021
-
[26]
Emerging properties in self-supervised vision transformers,
M. Caron, H. Touvron, I. Misra, H. J ´egou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” in Proceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 9650–9660
2021
-
[27]
ibot: Image bert pre-training with online tokenizer,
J. Zhou, C. Wei, H. Wang, W. Shen, C. Xie, A. Yuille, and T. Kong, “ibot: Image bert pre-training with online tokenizer,” arXiv preprint arXiv:2111.07832, 2021
2021 arXiv
-
[28]
Dinov2: Learning robust visual features without supervision,
M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby et al. , “Dinov2: Learning robust visual features without supervision,” arXiv preprint arXiv:2304.07193, 2023
2023 arXiv
-
[29]
Contrastive learning with stronger augmenta- tions,
X. Wang and G.-J. Qi, “Contrastive learning with stronger augmenta- tions,” IEEE transactions on pattern analysis and machine intelligence , vol. 45, no. 5, pp. 5549–5560, 2022
2022
-
[30]
What if we only use real datasets for scene text recognition? toward scene text recognition with fewer labels,
J. Baek, Y . Matsui, and K. Aizawa, “What if we only use real datasets for scene text recognition? toward scene text recognition with fewer labels,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2021, pp. 3113–3122
2021
-
[31]
Siman: Exploring self-supervised rep- resentation learning of scene text via similarity-aware normalization,
C. Luo, L. Jin, and J. Chen, “Siman: Exploring self-supervised rep- resentation learning of scene text via similarity-aware normalization,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 1039–1048
2022
-
[32]
Self- supervised character-to-character distillation for text recognition,
T. Guan, W. Shen, X. Yang, Q. Feng, Z. Jiang, and X. Yang, “Self- supervised character-to-character distillation for text recognition,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 19 473–19 484
2023
-
[33]
Reading scene text in deep convolutional sequences,
P. He, W. Huang, Y . Qiao, C. Loy, and X. Tang, “Reading scene text in deep convolutional sequences,” in Proceedings of the AAAI conference on artificial intelligence , vol. 30, no. 1, 2016
2016
-
[34]
Accurate recognition of words in scenes without char- acter segmentation using recurrent neural network,
B. Su and S. Lu, “Accurate recognition of words in scenes without char- acter segmentation using recurrent neural network,” Pattern Recognition, vol. 63, pp. 397–405, 2017
2017
-
[35]
An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition,
B. Shi, X. Bai, and C. Yao, “An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition,” IEEE transactions on pattern analysis and machine intelligence, vol. 39, no. 11, pp. 2298–2304, 2016
2016
-
[36]
Aster: An attentional scene text recognizer with flexible rectification,
B. Shi, M. Yang, X. Wang, P. Lyu, C. Yao, and X. Bai, “Aster: An attentional scene text recognizer with flexible rectification,” IEEE transactions on pattern analysis and machine intelligence, vol. 41, no. 9, pp. 2035–2048, 2018
2018
-
[37]
Learning to read irregular text with attention mechanisms
X. Yang, D. He, Z. Zhou, D. Kifer, and C. L. Giles, “Learning to read irregular text with attention mechanisms.” in IJCAI, vol. 1, no. 2, 2017, p. 3
2017
-
[38]
Attention-based extraction of structured information from street view imagery,
Z. Wojna, A. N. Gorban, D.-S. Lee, K. Murphy, Q. Yu, Y . Li, and J. Ibarz, “Attention-based extraction of structured information from street view imagery,” in2017 14th IAPR international conference on document analysis and recognition (ICDAR) , vol. 1. Ieee, 2017, pp. 844–850....
2017
-
[39]
On recognizing texts of arbitrary shapes with 2d self-attention,
J. Lee, S. Park, J. Baek, S. J. Oh, S. Kim, and H. Lee, “On recognizing texts of arbitrary shapes with 2d self-attention,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020, pp. 546–547
2020
-
[40]
Master: Multi-aspect non-local network for scene text recognition,
N. Lu, W. Yu, X. Qi, Y . Chen, P. Gong, R. Xiao, and X. Bai, “Master: Multi-aspect non-local network for scene text recognition,” Pattern Recognition, vol. 117, p. 107980, 2021
2021
-
[41]
Context-based contrastive learning for scene text recognition,
X. Zhang, B. Zhu, X. Yao, Q. Sun, R. Li, and B. Yu, “Context-based contrastive learning for scene text recognition,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 36, no. 3, 2022, pp. 3353–3361
2022
-
[42]
What do self-supervised vision transformers learn?
N. Park, W. Kim, B. Heo, T. Kim, and S. Yun, “What do self-supervised vision transformers learn?” arXiv preprint arXiv:2305.00729 , 2023
2023 arXiv
-
[43]
Masked siamese convnets,
L. Jing, J. Zhu, and Y . LeCun, “Masked siamese convnets,” arXiv preprint arXiv:2206.07700, 2022
2022 arXiv
-
[44]
Convnext v2: Co-designing and scaling convnets with masked autoen- coders,
S. Woo, S. Debnath, R. Hu, X. Chen, Z. Liu, I. S. Kweon, and S. Xie, “Convnext v2: Co-designing and scaling convnets with masked autoen- coders,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 16 133–16 142
2023
-
[45]
Scene text recognition using higher order language priors,
A. Mishra, K. Alahari, and C. Jawahar, “Scene text recognition using higher order language priors,” in BMVC-British machine vision confer- ence. BMV A, 2012
2012
-
[46]
Icdar 2003 robust reading competitions: entries, results, and future directions,
S. M. Lucas, A. Panaretos, L. Sosa, A. Tang, S. Wong, R. Young, K. Ashida, H. Nagai, M. Okamoto, H. Yamamoto et al. , “Icdar 2003 robust reading competitions: entries, results, and future directions,” International Journal of Document Analysis and Recognition (IJDAR) , vol. 7,...
2003
-
[47]
Icdar 2013 robust reading competition,
D. Karatzas, F. Shafait, S. Uchida, M. Iwamura, L. G. i Bigorda, S. R. Mestre, J. Mas, D. F. Mota, J. A. Almazan, and L. P. De Las Heras, “Icdar 2013 robust reading competition,” in 2013 12th international conference on document analysis and recognition . IEEE, 2013, pp. 1484–1493
2013
-
[48]
End-to-end scene text recog- nition,
K. Wang, B. Babenko, and S. Belongie, “End-to-end scene text recog- nition,” in 2011 International conference on computer vision . IEEE, 2011, pp. 1457–1464
2011
-
[49]
Icdar 2015 competition on robust reading,
D. Karatzas, L. Gomez-Bigorda, A. Nicolaou, S. Ghosh, A. Bagdanov, M. Iwamura, J. Matas, L. Neumann, V . R. Chandrasekhar, S. Lu et al., “Icdar 2015 competition on robust reading,” in 2015 13th international conference on document analysis and recognition (ICDAR) . IEEE, 2015,...
2015
-
[50]
Recognizing text with perspective distortion in natural scenes,
T. Q. Phan, P. Shivakumara, S. Tian, and C. L. Tan, “Recognizing text with perspective distortion in natural scenes,” inProceedings of the IEEE international conference on computer vision , 2013, pp. 569–576
2013
-
[51]
A robust arbitrary text detection system for natural scene images,
A. Risnumawan, P. Shivakumara, C. S. Chan, and C. L. Tan, “A robust arbitrary text detection system for natural scene images,” Expert Systems with Applications, vol. 41, no. 18, pp. 8027–8048, 2014
2014
-
[52]
Coco-text: Dataset and benchmark for text detection and recognition in natural images,
A. Veit, T. Matera, L. Neumann, J. Matas, and S. Belongie, “Coco-text: Dataset and benchmark for text detection and recognition in natural images,” arXiv preprint arXiv:1601.07140 , 2016
2016 arXiv
-
[53]
Curved scene text detection via transverse and longitudinal sequence connection,
Y . Liu, L. Jin, S. Zhang, C. Luo, and S. Zhang, “Curved scene text detection via transverse and longitudinal sequence connection,” Pattern Recognition, vol. 90, pp. 337–345, 2019
2019
-
[54]
Total-text: A comprehensive dataset for scene text detection and recognition,
C. K. Ch’ng and C. S. Chan, “Total-text: A comprehensive dataset for scene text detection and recognition,” in 2017 14th IAPR international conference on document analysis and recognition (ICDAR) , vol. 1. IEEE, 2017, pp. 935–942
2017
-
[55]
From two to one: A new scene text recognizer with visual language modeling network,
Y . Wang, H. Xie, S. Fang, J. Wang, S. Zhu, and Y . Zhang, “From two to one: A new scene text recognizer with visual language modeling network,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 14 194–14 203
2021
-
[56]
Robust scene text recognition with automatic rectification,
B. Shi, X. Wang, P. Lyu, C. Yao, and X. Bai, “Robust scene text recognition with automatic rectification,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 4168– 4176
2016
-
[57]
What is wrong with scene text recognition model comparisons? dataset and model analysis,
J. Baek, G. Kim, J. Lee, S. Park, D. Han, S. Yun, S. J. Oh, and H. Lee, “What is wrong with scene text recognition model comparisons? dataset and model analysis,” in Proceedings of the IEEE/CVF international conference on computer vision , 2019, pp. 4715–4723
2019
-
[58]
Synthetic data and artificial neural networks for natural scene text recognition,
M. Jaderberg, K. Simonyan, A. Vedaldi, and A. Zisserman, “Synthetic data and artificial neural networks for natural scene text recognition,” arXiv preprint arXiv:1406.2227 , 2014
2014 arXiv
-
[59]
Benchmarking chinese text recognition: Datasets, baselines, and an empirical study,
H. Yu, J. Chen, B. Li, J. Ma, M. Guan, X. Xu, X. Wang, S. Qu, and X. Xue, “Benchmarking chinese text recognition: Datasets, baselines, and an empirical study,” arXiv preprint arXiv:2112.15093 , 2021
2021 arXiv
-
[60]
The iam-database: an english sentence database for offline handwriting recognition,
U.-V . Marti and H. Bunke, “The iam-database: an english sentence database for offline handwriting recognition,” International Journal on Document Analysis and Recognition , vol. 5, pp. 39–46, 2002
2002
-
[61]
Cvl-database: An off-line database for writer retrieval, writer identification and word spotting,
F. Kleber, S. Fiel, M. Diem, and R. Sablatnig, “Cvl-database: An off-line database for writer retrieval, writer identification and word spotting,” in 2013 12th international conference on document analysis and recognition, IEEE. IEEE, 2013, pp. 560–564
2013
-
[62]
Stochastic neighbor embedding,
G. E. Hinton and S. Roweis, “Stochastic neighbor embedding,” Advances in neural information processing systems , vol. 15, 2002
2002
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.