Fetching the paper…
Reading the bibliography…
Due to the enormous technical challenges and wide range of applications, scene text recognition (STR) has been an active research topic in computer vision for years.
U. Marti and H. Bunke, “The iam-database: an english sentence database for offline handwriting recognition,” Int. J. Document Anal. Recognit. , vol. 5, no. 1, pp. 39–46, 2002
2002
Earlier work this paper cites.
A. Graves, S. Fernández, F. J. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Int. Conf. Mach. Learn. , vol. 148, 2006, pp. 369–376
2006
Earlier work this paper cites.
A. Yuille and D. Kersten, “Vision as bayesian inference: analysis by synthesis?” Trends in cognitive sciences , vol. 10, no. 7, pp. 301–308, 2006
2006
Earlier work this paper cites.
E. Grosicki, M. Carré, J. Brodin, and E. Geoffrois, “Results of the RIMES evaluation campaign for handwritten mail processing,” in 10th International Conference on Document Analysis and Recognition, ICDAR 2009, Barcelona, Spain, 26-29 July 2009 . IEEE Computer Society, 2009, pp. 941–945
2009
Earlier work this paper cites.
K. Wang, B. Babenko, and S. J. Belongie, “End-to-end scene text recognition,” in Int. Conf. Comput. Vis. , 2011, pp. 1457–1464
2011
Earlier work this paper cites.
A. Mishra, K. Alahari, and C. V. Jawahar, “Scene text recognition using higher order language priors,” in BMVC , 2012, pp. 1–11
2012
Earlier work this paper cites.
M. Schuster and K. Nakajima, “Japanese and korean voice search,” in ICASSP , 2012, pp. 5149–5152
2012
Earlier work this paper cites.
M. D. Zeiler, “ADADELTA: an adaptive learning rate method,” CoRR , vol. abs/1212.5701, 2012
2012
Earlier work this paper cites.
D. Karatzas, F. Shafait, S. Uchida, M. Iwamura, L. G. i Bigorda, S. R. Mestre, J. Mas, D. F. Mota, J. Almazán, and L. de las Heras, “ICDAR 2013 robust reading competition,” in ICDAR , 2013, pp. 1484–1493
2013
Earlier work this paper cites.
T. Q. Phan, P. Shivakumara, S. Tian, and C. L. Tan, “Recognizing text with perspective distortion in natural scenes,” in Int. Conf. Comput. Vis. , 2013, pp. 569–576
2013
Earlier work this paper cites.
F. Kleber, S. Fiel, M. Diem, and R. Sablatnig, “Cvl-database: An off-line database for writer retrieval, writer identification and word spotting,” in ICDAR , 2013, pp. 560–564
2013
Earlier work this paper cites.
A. Risnumawan, P. Shivakumara, C. S. Chan, and C. L. Tan, “A robust arbitrary text detection system for natural scene images,” Expert Syst. Appl. , vol. 41, no. 18, pp. 8027–8048, 2014
2014
Earlier work this paper cites.
I. J. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. C. Courville, and Y. Bengio, “Generative adversarial nets,” in Adv. Neural Inform. Process. Syst. , 2014, pp. 2672–2680
2014
Earlier work this paper cites.
M. Jaderberg, K. Simonyan, A. Vedaldi, and A. Zisserman, “Synthetic data and artificial neural networks for natural scene text recognition,” NIPS Deep Learning Workshop , 2014
2014
Earlier work this paper cites.
D. Karatzas, L. Gomez-Bigorda, A. Nicolaou, S. K. Ghosh, A. D. Bagdanov, M. Iwamura, J. Matas, L. Neumann, V. R. Chandrasekhar, S. Lu, F. Shafait, S. Uchida, and E. Valveny, “ICDAR 2015 competition on robust reading,” in ICDAR , 2015, pp. 1156–1160
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in Int. Conf. Learn. Represent. , 2015
2015
Earlier work this paper cites.
Y. Zhu, C. Yao, and X. Bai, “Scene text detection and recognition: Recent advances and future trends,” Frontiers of Computer Science , vol. 10, no. 1, pp. 19–36, 2016
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Priya, S. Mishra, S. Raj, S. Mandal, and S. Datta, “Online and offline character recognition: A survey,” in 2016 International conference on communication and signal processing (ICCSP) . IEEE, 2016, pp. 0967–0970
2016
Earlier work this paper cites.
A. Purohit and S. S. Chauhan, “A literature survey on handwritten character recognition,” International Journal of Computer Science and Information Technologies (IJCSIT) , vol. 7, no. 1, pp. 1–5, 2016
2016
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2016, pp. 770–778
2016
Earlier work this paper cites.
C. Lee and S. Osindero, “Recursive recurrent nets with attention modeling for OCR in the wild,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2016, pp. 2231–2239
2016
Earlier work this paper cites.
B. Shi, X. Wang, P. Lyu, C. Yao, and X. Bai, “Robust scene text recognition with automatic rectification,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2016, pp. 4168–4176
2016
Earlier work this paper cites.
W. Liu, C. Chen, K. K. Wong, Z. Su, and J. Han, “Star-net: A spatial attention residue network for scene text recognition,” in BMVC , 2016
2016
Earlier work this paper cites.
P. He, W. Huang, Y. Qiao, C. C. Loy, and X. Tang, “Reading scene text in deep convolutional sequences,” in AAAI , 2016, pp. 3501–3508
2016
Earlier work this paper cites.
A. Poznanski and L. Wolf, “Cnn-n-gram for handwriting word recognition,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2016, pp. 2305–2314
2016
Earlier work this paper cites.
S. Sudholt and G. A. Fink, “Phocnet: A deep convolutional neural network for word spotting in handwritten documents,” in 2016 15th International Conference on Frontiers in Handwriting Recognition (ICFHR) . IEEE, 2016, pp. 277–282
2016
Earlier work this paper cites.
R. Sennrich, B. Haddow, and A. Birch, “Neural machine translation of rare words with subword units,” in ACL . The Association for Computer Linguistics, 2016
2016
Earlier work this paper cites.
——, “Reading text in the wild with convolutional neural networks,” Int. J. Comput. Vis. , vol. 116, no. 1, pp. 1–20, 2016
2016
Earlier work this paper cites.
A. Gupta, A. Vedaldi, and A. Zisserman, “Synthetic data for text localisation in natural images,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2016, pp. 2315–2324
2016
Earlier work this paper cites.
B. Shi, X. Bai, and C. Yao, “An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 39, no. 11, pp. 2298–2304, 2017
2017
Earlier work this paper cites.
Z. Cheng, F. Bai, Y. Xu, G. Zheng, S. Pu, and S. Zhou, “Focusing attention: Towards accurate text recognition in natural images,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2017, pp. 5086–5094
2017
Earlier work this paper cites.
Y. Zhang, L. Gueguen, I. Zharkov, P. Zhang, K. Seifert, and B. Kadlec, “Uber-text: A large-scale dataset for optical character recognition from street-level imagery,” in SUNw: Scene Understanding Workshop-CVPR , vol. 2017, 2017, pp. 1–5
2017
Earlier work this paper cites.
J. Wang and X. Hu, “Gated recurrent convolution neural network for OCR,” in Adv. Neural Inform. Process. Syst. , 2017, pp. 335–344
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Adv. Neural Inform. Process. Syst. , 2017, pp. 5998–6008
2017
Earlier work this paper cites.
M. Labeau and A. Allauzen, “Character and subword-based word representation for neural language modeling prediction,” in SWCN@EMNLP , 2017, pp. 1–13
2017
Earlier work this paper cites.
B. Shi, C. Yao, M. Liao, M. Yang, P. Xu, L. Cui, S. J. Belongie, S. Lu, and X. Bai, “ICDAR2017 competition on reading chinese text in the wild (RCTW-17),” in ICDAR , 2017, pp. 1429–1434
2017
Cited alongside, same era.
I. Loshchilov and F. Hutter, “SGDR: stochastic gradient descent with warm restarts,” in Int. Conf. Learn. Represent. , 2017
2017
Cited alongside, same era.
F. Borisyuk, A. Gordo, and V. Sivakumar, “Rosetta: Large scale system for text detection and recognition in images,” in SIGKDD , Y. Guo and F. Farooq, Eds., 2018, pp. 71–79
2018
Cited alongside, same era.
J. Sueiras, V. Ruiz, A. Sanchez, and J. F. Velez, “Offline continuous handwriting recognition using sequence to sequence neural networks,” Neurocomputing , vol. 289, pp. 119–128, 2018
2018
Cited alongside, same era.
X. Chen, L. Jin, Y. Zhu, C. Luo, and T. Wang, “Text recognition in the wild: A survey,” ACM Computing Surveys (CSUR) , vol. 54, no. 2, pp. 1–35, 2021
2021
Later among the works it cites.
N. Lu, W. Yu, X. Qi, Y. Chen, P. Gong, R. Xiao, and X. Bai, “MASTER: Multi-aspect non-local network for scene text recognition,” Pattern Recognition , vol. 117, p. 107980, 2021
2021
Later among the works it cites.
S. Fang, H. Xie, Y. Wang, Z. Mao, and Y. Zhang, “Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2021, pp. 7098–7107
2021
Later among the works it cites.
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in Int. Conf. Learn. Represent. , 2021
2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
2018
Cited alongside, same era.
B. Shi, M. Yang, X. Wang, P. Lyu, C. Yao, and X. Bai, “ASTER: an attentional scene text recognizer with flexible rectification,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 41, no. 9, pp. 2035–2048, 2019
2019
Cited alongside, same era.
C. K. Chng, E. Ding, J. Liu, D. Karatzas, C. S. Chan, L. Jin, Y. Liu, Y. Sun, C. C. Ng, C. Luo, Z. Ni, C. Fang, S. Zhang, and J. Han, “ICDAR2019 robust reading challenge on arbitrary-shaped text - rrc-art,” in ICDAR , 2019, pp. 1571–1576
2019
Cited alongside, same era.
J. Baek, G. Kim, J. Lee, S. Park, D. Han, S. Yun, S. J. Oh, and H. Lee, “What is wrong with scene text recognition model comparisons? dataset and model analysis,” in Int. Conf. Comput. Vis. , 2019, pp. 4714–4722
2019
Cited alongside, same era.
2019
Cited alongside, same era.
M. Liao, J. Zhang, Z. Wan, F. Xie, J. Liang, P. Lyu, C. Yao, and X. Bai, “Scene text recognition from two-dimensional perspective,” in AAAI , 2019, pp. 8714–8721
2019
Cited alongside, same era.
A. K. Bhunia, A. Das, A. K. Bhunia, P. S. R. Kishore, and P. P. Roy, “Handwriting recognition in low-resource scripts using adversarial learning,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2019, pp. 4767–4776
2019
Cited alongside, same era.
2021
Later among the works it cites.
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervision,” in Int. Conf. Mach. Learn. , vol. 139, 2021, pp. 8748–8763
2021
Later among the works it cites.
M. Liao, P. Lyu, M. He, C. Yao, W. Wu, and X. Bai, “Mask textspotter: An end-to-end trainable neural network for spotting text with arbitrary shapes,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 43, no. 2, pp. 532–548, 2021
2021
Later among the works it cites.
R. Atienza, “Vision transformer for fast and efficient scene text recognition,” in ICDAR , vol. 12821, 2021, pp. 319–334
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Int. Conf. Comput. Vis. , 2021, pp. 9992–10 002
2021
Later among the works it cites.
2021
Later among the works it cites.
Y. Wang, H. Xie, S. Fang, J. Wang, S. Zhu, and Y. Zhang, “From two to one: A new scene text recognizer with visual language modeling network,” in Int. Conf. Comput. Vis. , 2021, pp. 1–10
2021
Later among the works it cites.
A. Aberdam, R. Litman, S. Tsiper, O. Anschel, R. Slossberg, S. Mazor, R. Manmatha, and P. Perona, “Sequence-to-sequence contrastive learning for text recognition,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2021, pp. 15 302–15 312
2021
Later among the works it cites.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in Int. Conf. Mach. Learn. , vol. 139, 2021, pp. 10 347–10 357
2021
Later among the works it cites.
A. Singh, G. Pang, M. Toh, J. Huang, W. Galuba, and T. Hassner, “Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2021, pp. 8802–8812
2021
Later among the works it cites.
I. Krylov, S. Nosov, and V. Sovrasov, “Open images V5 text annotation and yet another mask text spotter,” in Asian Conference on Machine Learning , V. N. Balasubramanian and I. W. Tsang, Eds., vol. 157, 2021, pp. 379–389
2021
Later among the works it cites.
J. Baek, Y. Matsui, and K. Aizawa, “What if we only use real datasets for scene text recognition? toward scene text recognition with fewer labels,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2021, pp. 3113–3122
2021
Later among the works it cites.
M. Caron, H. Touvron, I. Misra, H. Jégou, J. Mairal, P. Bojanowski, and A. Joulin, “Emerging properties in self-supervised vision transformers,” in Int. Conf. Comput. Vis. , 2021, pp. 9630–9640
2021
Later among the works it cites.
B. Na, Y. Kim, and S. Park, “Multi-modal text recognition networks: Interactive enhancements between visual and semantic features,” in Eur. Conf. Comput. Vis. , vol. 13688, 2022, pp. 446–463
2022
Later among the works it cites.
P. Wang, C. Da, and C. Yao, “Multi-granularity prediction for scene text recognition,” in Eur. Conf. Comput. Vis. , vol. 13688, 2022, pp. 339–355
2022
Later among the works it cites.
X. Zhang, B. Zhu, X. Yao, Q. Sun, R. Li, and B. Yu, “Context-based contrastive learning for scene text recognition,” in AAAI , 2022, pp. 888–896
2022
Later among the works it cites.
H. Liu, B. Wang, Z. Bao, M. Xue, S. Kang, D. Jiang, Y. Liu, and B. Ren, “Perceiving stroke-semantic context: Hierarchical contrastive learning for robust scene text recognition,” in AAAI , 2022, pp. 1702–1710
2022
Later among the works it cites.
Y. L. Tan, A. W. Kong, and J. Kim, “Pure transformer with integrated experts for scene text recognition,” in Eur. Conf. Comput. Vis. , vol. 13688, 2022, pp. 481–497
2022
Later among the works it cites.
C. Da, P. Wang, and C. Yao, “Levenshtein OCR,” in Eur. Conf. Comput. Vis. , vol. 13688, 2022, pp. 322–338
2022
Later among the works it cites.
D. Bautista and R. Atienza, “Scene text recognition with permuted autoregressive sequence models,” in Eur. Conf. Comput. Vis. , vol. 13688, 2022, pp. 178–196
2022
Later among the works it cites.
2022
Later among the works it cites.
M. Yang, M. Liao, P. Lu, J. Wang, S. Zhu, H. Luo, Q. Tian, and X. Bai, “Reading and writing: Discriminative and generative modeling for self-supervised text recognition,” in MM ’22: The 30th ACM International Conference on Multimedia, Lisboa, Portugal, October 10 - 14, 2022 . ACM, 2022, pp. 4214–4223
2022
Later among the works it cites.
K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. B. Girshick, “Masked autoencoders are scalable vision learners,” in IEEE Conf. Comput. Vis. Pattern Recog. , 2022, pp. 15 979–15 988
2022
Later among the works it cites.
J. Li, D. Li, C. Xiong, and S. C. H. Hoi, “BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation,” in Int. Conf. Mach. Learn. , vol. 162, 2022, pp. 12 888–12 900
2022
Later among the works it cites.
O. Nuriel, S. Fogel, and R. Litman, “Textadain: Paying attention to shortcut learning in text recognizers,” in Eur. Conf. Comput. Vis. , vol. 13688, 2022, pp. 427–445
2022
Later among the works it cites.
F.-L. Chen, D.-Z. Zhang, M.-L. Han, X.-Y. Chen, J. Shi, S. Xu, and B. Xu, “Vlp: A survey on vision-language pre-training,” Machine Intelligence Research , vol. 20, no. 1, pp. 38–56, 2023
2023
Closest in time.
M. Li, T. Lv, L. Cui, Y. Lu, D. A. F. Florêncio, C. Zhang, Z. Li, and F. Wei, “Trocr: Transformer-based optical character recognition with pre-trained models,” in AAAI , 2023, pp. 13 094–13 102
2023
Closest in time.
2023
Closest in time.
S. Fang, Z. Mao, H. Xie, Y. Wang, C. Yan, and Y. Zhang, “Abinet++: Autonomous, bidirectional and iterative language modeling for scene text spotting,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 45, no. 6, pp. 7123–7141, 2023
2023
Closest in time.
Z. Qiao, Y. Zhou, J. Wei, W. Wang, Y. Zhang, N. Jiang, H. Wang, and W. Wang, “Pimnet: A parallel, iterative and mimicking network for scene text recognition,” in ACM Int. Conf. Multimedia , 2021, pp. 2046–2055
2055
Closest in time.
F. Zhan and S. Lu, “ESIR: end-to-end scene text recognition via iterative image rectification,” in CVPR , 2019, pp. 2059–2068
2068
Closest in time.