Fetching the paper…
Reading the bibliography…
Scene text recognition (STR) methods have struggled to attain high accuracy and fast inference speed.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Comput. , pp. 1735–1780, 1997
1997
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,” in ICML , 2006, p. 369–376
2006
Earlier work this paper cites.
K. Wang, B. Babenko, and S. Belongie, “End-to-end scene text recognition,” in ICCV , 2011, pp. 1457–1464
2011
Earlier work this paper cites.
A. Mishra, A. Karteek, and C. V. Jawahar, “Scene text recognition using higher order language priors,” in BMVC , 2012
2012
Earlier work this paper cites.
D. KaratzasAU, F. ShafaitAU, S. UchidaAU, M. IwamuraAU, L. G. i. BigordaAU, S. R. MestreAU, J. MasAU, D. F. MotaAU, J. A. AlmazànAU, and L. P. de las Heras, “Icdar 2013 robust reading competition,” in ICDAR , 2013, pp. 1484–1493
2013
Earlier work this paper cites.
T. Q. Phan, P. Shivakumara, S. Tian, and C. L. Tan, “Recognizing text with perspective distortion in natural scenes,” in CVPR , 2013, pp. 569–576
2013
Earlier work this paper cites.
I. Sutskever, O. Vinyals, and Q. V. Le, “Sequence to sequence learning with neural networks,” in NIPS , 2014, pp. 3104–3112
2014
Earlier work this paper cites.
K. Cho, B. Merrienboer, Ç. Gülçehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using rnn encoder-decoder for statistical machine translation,” in EMNLP , 2014, pp. 1724–1734
2014
Earlier work this paper cites.
J. Chung, Ç. Gülçehre, K. Cho, and Y. Bengio, “Empirical evaluation of gated recurrent neural networks on sequence modeling,” 2014
2014
Earlier work this paper cites.
M. Jaderberg, K. Simonyan, A. Vedaldi, and A. Zisserman, “Synthetic data and artificial neural networks for natural scene text recognition,” in NeurIPS Deep Learning Workshop , 2014
2014
Earlier work this paper cites.
R. Anhar, S. Palaiahnakote, C. S. Chan, and C. L. Tan, “A robust arbitrary text detection system for natural scene images,” in Expert Systems with Applications , 2014, p. 8027–8048
2014
Earlier work this paper cites.
M. Jaderberg, K. Simonyan, A. Vedaldi, and A. Zisserman, “Reading text in the wild with convolutional neural networks,” in IJCV , 2015, pp. 1–20
2015
Earlier work this paper cites.
D. Karatzas, L. Gomez-Bigorda, A. Nicolaou, S. Ghosh, A. Bagdanov, M. Iwamura, J. Matas, L. Neumann, V. R. Chandrasekhar, S. Lu, F. Shafait, S. Uchida, and E. Valveny, “Icdar 2015 competition on robust reading,” in ICDAR , 2015, pp. 1156–1160
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and S., “Deep residual learning for image recognition,” in CVPR , June 2016
2016
Earlier work this paper cites.
A. Gupta, A. Vedaldi, and A. Zisserman, “Synthetic data for text localisation in natural images,” in CVPR , 2016, pp. 2315–2324
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in NIPS , 2017
2017
Earlier work this paper cites.
B. Shi, X. Bai, and C. Yao, “An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition,” in TPAMI , 2017, pp. 2298–2304
2017
Earlier work this paper cites.
I. Loshchilov and F. Hutter, “SGDR: stochastic gradient descent with warm restarts,” in ICLR , 2017
2017
Earlier work this paper cites.
H. Li, P. Wang, C. Shen, and G. Zhang, “Show, attend and read: A simple and strong baseline for irregular text recognition,” in AAAI , 2019, pp. 8610–8617
2019
Earlier work this paper cites.
F. Sheng, Z. Chen, and B. Xu, “Nrtr: A no-recurrence sequence-to-sequence model for scene text recognition,” in ICDAR , 2019, pp. 781–786
2019
Earlier work this paper cites.
Z. Xie, Y. Huang, Y. Zhu, L. Jin, Y. Liu, and L. Xie, “Aggregation cross-entropy for sequence recognition,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , 2019, pp. 6538–6547
2019
Cited alongside, same era.
J. Baek, G. Kim, J. Lee, S. Park, D. Han, S. Yun, S. J. Oh, and H. Lee, “What is wrong with scene text recognition model comparisons? dataset and model analysis,” in ICCV , 2019, pp. 4714–4722
2019
Cited alongside, same era.
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” in ICLR , 2019
2019
Cited alongside, same era.
C. Luo, L. Jin, and Z. Sun, “Moran: A multi-object rectified attention network for scene text recognition,” Pattern Recogn. , p. 109–118, 2019
2019
Cited alongside, same era.
Z. Wan, M. He, H. Chen, X. Bai, and C. Yao, “Textscanner: Reading characters in order for robust scene text recognition,” in AAAI , vol. 34, no. 07, 2020, pp. 12 120–12 127
J. Chen, B. Li, and X. Xue, “Scene text telescope: Text-focused scene image super-resolution,” in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2021, pp. 12 021–12 030
2021
Later among the works it cites.
2021
Later among the works it cites.
D. Bautista and R. Atienza, “Scene text recognition with permuted autoregressive sequence models,” in ECCV , 2022, pp. 178–196
2022
Later among the works it cites.
Y. Du, Z. Chen, C. Jia, X. Yin, T. Zheng, C. Li, Y. Du, and Y. Jiang, “Svtr: Scene text recognition with a single visual model,” in IJCAI , 2022, pp. 884–890
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
X. Yue, Z. Kuang, C. Lin, H. Sun, and W. Zhang, “Robustscanner: Dynamically enhancing positional clues for robust text recognition,” in ECCV , 2020, pp. 135–151
2020
Cited alongside, same era.
D. Yu, X. Li, C. Zhang, T. Liu, J. Han, J. Liu, and E. Ding, “Towards accurate scene text recognition with semantic reasoning networks,” in CVPR , 2020, pp. 12 113–12 122
2020
Cited alongside, same era.
W. Hu, X. Cai, J. Hou, S. Yi, and Z. Lin, “Gtc: Guided training of ctc towards efficient and accurate scene text recognition,” in AAAI , 2020, pp. 11 005–11 012
2020
Cited alongside, same era.
T. Wang, Y. Zhu, L. Jin, C. Luo, X. Chen, Y. Wu, Q. Wang, and M. Cai, “Decoupled attention network for text recognition,” in Proceedings of the AAAI conference on artificial intelligence , vol. 34, no. 07, 2020, pp. 12 216–12 224
2020
Cited alongside, same era.
J. Lee, S. Park, J. Baek, S. Oh, S. Kim, and H. Lee, “On recognizing texts of arbitrary shapes with 2d self-attention,” in CVPR Workshops , 2020, pp. 546–547
2020
Cited alongside, same era.
Z. Qiao, Y. Zhou, D. Yang, Y. Zhou, and W. Wang, “Seed: Semantics enhanced encoder-decoder framework for scene text recognition,” in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , 2020, pp. 13 525–13 534
2020
Cited alongside, same era.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in ICLR , 2021
2021
Cited alongside, same era.
2022
Later among the works it cites.
X. Xie, L. Fu, Z. Zhang, Z. Wang, and X. Bai, “Toward understanding wordart: Corner-guided transformer for scene text recognition,” in ECCV . Springer, 2022, pp. 303–321
2022
Later among the works it cites.
B. Li, Y. Yuan, D. Liang, X. Liu, Z. Ji, J. Bai, W. Liu, and X. Bai, “When counting meets hmer: Counting-aware network for handwritten mathematical expression recognition,” in ECCV . Springer, 2022, pp. 197–214
2022
Later among the works it cites.
2022
Later among the works it cites.
Y. He, C. Chen, J. Zhang, J. Liu, F. He, C. Wang, and B. Du, “Visual semantics allow for textual reasoning better in scene text recognition,” in AAAI , 2022, pp. 888–896
2022
Later among the works it cites.
T. Zheng, Z. Chen, S. Fang, H. Xie, and Y.-G. Jiang, “Cdistnet: Perceiving multi-domain character distance for robust text recognition,” International Journal of Computer Vision , pp. 1–19, 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
S. Fang, Z. Mao, H. Xie, Y. Wang, C. Yan, and Y. Zhang, “Abinet++: Autonomous, bidirectional and iterative language modeling for scene text spotting,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 45, no. 6, p. 7123–7141, 2023
2023
Closest in time.
C. Xue, J. Huang, W. Zhang, S. Lu, C. Wang, and S. Bai, “Image-to-character-to-word transformers for accurate scene text recognition,” IEEE Trans. Pattern Anal. Mach. Intell. , 2023
2023
Closest in time.
B. Zhang, H. Xie, Y. Wang, J. Xu, and Y. Zhang, “Linguistic more: Taking a further step toward efficient and accurate scene text recognition.” International Joint Conferences on Artificial Intelligence Organization, 2023, pp. 1704–1712
2023
Closest in time.
H. Yu, X. Wang, B. Li, and X. Xue, “Chinese text recognition with a pre-trained clip-like model through image-ids aligning,” in ICCV , 2023
2023
Closest in time.
M. Awais, M. Naseer, S. Khan, R. M. Anwer, H. Cholakkal, M. Shah, M.-H. Yang, and F. S. Khan, “Foundational models defining a new era in vision: A survey and outlook,” 2023
2023
Closest in time.
B. Shi, M. Yang, X. Wang, P. Lyu, C. Yao, and X. Bai, “Aster: An attentional scene text recognizer with flexible rectification,” in IEEE Trans. Pattern Anal. Mach. Intell. , 2019, pp. 2035–2048
2048
Closest in time.
Z. Qiao, Y. Zhou, J. Wei, W. Wang, Y. Zhang, N. Jiang, H. Wang, and W. Wang, “Pimnet: a parallel, iterative and mimicking network for scene text recognition,” in Proceedings of the 29th ACM International Conference on Multimedia , 2021, pp. 2046–2055
2055
Closest in time.