Fetching the paper…
Reading the bibliography…
Leveraging the advances of natural language processing, most recent scene text recognizers adopt an encoder-decoder architecture where text images are first converted to representative features and then a sequence of characters via `sequential decoding'.
F. L. Bookstein, “Principal warps: Thin-plate splines and the decomposition of deformations,” IEEE Transactions on pattern analysis and machine intelligence , vol. 11, no. 6, pp. 567–585, 1989
1989
Earlier work this paper cites.
J. L. Elman, “Finding structure in time,” Cognitive science , vol. 14, no. 2, pp. 179–211, 1990
1990
Earlier work this paper cites.
S. M. Lucas, A. Panaretos, L. Sosa, A. Tang, S. Wong, R. Young, K. Ashida, H. Nagai, M. Okamoto, H. Yamamoto et al. , “Icdar 2003 robust reading competitions: entries, results, and future directions,” International Journal of Document Analysis and Recognition (IJDAR) , vol. 7, no. 2-3, pp. 105–122, 2005
2005
Earlier work this paper cites.
A. Graves, S. Fernández, F. Gomez, and J. Schmidhuber, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” in Proceedings of the 23rd international conference on Machine learning , 2006, pp. 369–376
2006
Earlier work this paper cites.
K. Wang, B. Babenko, and S. Belongie, “End-to-end scene text recognition,” in 2011 International Conference on Computer Vision . IEEE, 2011, pp. 1457–1464
2011
Earlier work this paper cites.
A. Mishra, K. Alahari, and C. Jawahar, “Top-down and bottom-up cues for scene text recognition,” in 2012 IEEE Conference on Computer Vision and Pattern Recognition . IEEE, 2012, pp. 2687–2694
2012
Earlier work this paper cites.
T. Q. Phan, P. Shivakumara, S. Tian, and C. L. Tan, “Recognizing text with perspective distortion in natural scenes,” in Proceedings of the IEEE International Conference on Computer Vision , 2013, pp. 569–576
2013
Earlier work this paper cites.
D. Karatzas, F. Shafait, S. Uchida, M. Iwamura, L. G. i Bigorda, S. R. Mestre, J. Mas, D. F. Mota, J. A. Almazan, and L. P. De Las Heras, “Icdar 2013 robust reading competition,” in 2013 12th International Conference on Document Analysis and Recognition . IEEE, 2013, pp. 1484–1493
2013
Earlier work this paper cites.
2014
Earlier work this paper cites.
B. Su and S. Lu, “Accurate scene text recognition based on recurrent neural network,” in Asian Conference on Computer Vision . Springer, 2014, pp. 35–48
2014
Earlier work this paper cites.
2014
Earlier work this paper cites.
A. Risnumawan, P. Shivakumara, C. S. Chan, and C. L. Tan, “A robust arbitrary text detection system for natural scene images,” Expert Systems with Applications , vol. 41, no. 18, pp. 8027–8048, 2014
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
D. Karatzas, L. Gomez-Bigorda, A. Nicolaou, S. Ghosh, A. Bagdanov, M. Iwamura, J. Matas, L. Neumann, V. R. Chandrasekhar, S. Lu et al. , “Icdar 2015 competition on robust reading,” in 2015 13th International Conference on Document Analysis and Recognition (ICDAR) . IEEE, 2015, pp. 1156–1160
2015
Earlier work this paper cites.
A. Gupta, A. Vedaldi, and A. Zisserman, “Synthetic data for text localisation in natural images,” in IEEE Conference on Computer Vision and Pattern Recognition , 2016
2016
Earlier work this paper cites.
M. Jaderberg, K. Simonyan, A. Vedaldi, and A. Zisserman, “Reading text in the wild with convolutional neural networks,” International journal of computer vision , vol. 116, no. 1, pp. 1–20, 2016
2016
Earlier work this paper cites.
P. He, W. Huang, Y. Qiao, C. C. Loy, and X. Tang, “Reading scene text in deep convolutional sequences,” in Thirtieth AAAI conference on artificial intelligence , 2016
2016
Earlier work this paper cites.
B. Shi, X. Bai, and C. Yao, “An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition,” IEEE transactions on pattern analysis and machine intelligence , vol. 39, no. 11, pp. 2298–2304, 2016
2016
Earlier work this paper cites.
C.-Y. Lee and S. Osindero, “Recursive recurrent nets with attention modeling for ocr in the wild,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2016, pp. 2231–2239
2016
Earlier work this paper cites.
W. Liu, C. Chen, K.-Y. K. Wong, Z. Su, and J. Han, “Star-net: a spatial attention residue network for scene text recognition.” in BMVC , vol. 2, 2016, p. 7
2016
Earlier work this paper cites.
B. Shi, X. Wang, P. Lyu, C. Yao, and X. Bai, “Robust scene text recognition with automatic rectification,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 4168–4176
2016
Earlier work this paper cites.
A. Gupta, A. Vedaldi, and A. Zisserman, “Synthetic data for text localisation in natural images,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2016, pp. 2315–2324
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
X. Yang, D. He, Z. Zhou, D. Kifer, and C. L. Giles, “Learning to read irregular text with attention mechanisms.” in IJCAI , vol. 1, no. 2, 2017, p. 3
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
C. K. Ch’ng and C. S. Chan, “Total-text: A comprehensive dataset for scene text detection and recognition,” in 2017 14th IAPR International Conference on Document Analysis and Recognition (ICDAR) , vol. 1. IEEE, 2017, pp. 935–942
2017
Earlier work this paper cites.
F. Zhan, S. Lu, and C. Xue, “Verisimilar image synthesis for accurate detection and recognition of texts in scenes,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 249–266
2018
Earlier work this paper cites.
B. Shi, M. Yang, X. Wang, P. Lyu, C. Yao, and X. Bai, “Aster: An attentional scene text recognizer with flexible rectification,” IEEE transactions on pattern analysis and machine intelligence , vol. 41, no. 9, pp. 2035–2048, 2018
2018
Cited alongside, same era.
F. Bai, Z. Cheng, Y. Niu, S. Pu, and S. Zhou, “Edit probability for scene text recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 1508–1516
2018
Cited alongside, same era.
Z. Liu, Y. Li, F. Ren, W. L. Goh, and H. Yu, “Squeezedtext: A real-time scene text recognition by binary convolutional encoder-decoder network,” in Thirty-Second AAAI Conference on Artificial Intelligence , 2018
2018
Cited alongside, same era.
Z. Cheng, Y. Xu, F. Bai, Y. Niu, S. Pu, and S. Zhou, “Aon: Towards arbitrarily-oriented text recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , 2018, pp. 5571–5579
2018
Cited alongside, same era.
W. Wang, E. Xie, X. Liu, W. Wang, D. Liang, C. Shen, and X. Bai, “Scene text image super-resolution in the wild,” in European Conference on Computer Vision . Springer, 2020, pp. 650–666
2020
Later among the works it cites.
Z. Wan, M. He, H. Chen, X. Bai, and C. Yao, “Textscanner: Reading characters in order for robust scene text recognition,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 07, 2020, pp. 12 120–12 127
2020
Later among the works it cites.
S. Long, Y. Guan, K. Bian, and C. Yao, “A new perspective for flexible feature gathering in scene text recognition via character anchor pooling,” in ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2020, pp. 2458–2462
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Lyu, M. Liao, C. Yao, W. Wu, and X. Bai, “Mask textspotter: An end-to-end trainable neural network for spotting text with arbitrary shapes,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 67–83
2018
Cited alongside, same era.
W. Liu, C. Chen, and K.-Y. K. Wong, “Char-net: A character-aware neural network for distorted scene text recognition,” in Thirty-Second AAAI Conference on Artificial Intelligence , 2018
2018
Cited alongside, same era.
2018
Cited alongside, same era.
N. Parmar, A. Vaswani, J. Uszkoreit, L. Kaiser, N. Shazeer, A. Ku, and D. Tran, “Image transformer,” in International Conference on Machine Learning . PMLR, 2018, pp. 4055–4064
2018
Cited alongside, same era.
2018
Cited alongside, same era.
F. Zhan, H. Zhu, and S. Lu, “Spatial fusion gan for image synthesis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 3653–3662
2019
Cited alongside, same era.
Z. Xie, Y. Huang, Y. Zhu, L. Jin, Y. Liu, and L. Xie, “Aggregation cross-entropy for sequence recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 6538–6547
2019
Cited alongside, same era.
M. Yang, Y. Guan, M. Liao, X. He, K. Bian, S. Bai, C. Yao, and X. Bai, “Symmetry-constrained rectification network for scene text recognition,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2019, pp. 9147–9156
2019
Cited alongside, same era.
2020
Later among the works it cites.
N. Carion, F. Massa, G. Synnaeve, N. Usunier, A. Kirillov, and S. Zagoruyko, “End-to-end object detection with transformers,” in European Conference on Computer Vision . Springer, 2020, pp. 213–229
2020
Later among the works it cites.
2020
Later among the works it cites.
M. Chen, A. Radford, R. Child, J. Wu, H. Jun, D. Luan, and I. Sutskever, “Generative pretraining from pixels,” in International Conference on Machine Learning . PMLR, 2020, pp. 1691–1703
2020
Later among the works it cites.
F. Yang, H. Yang, J. Fu, H. Lu, and B. Guo, “Learning texture transformer network for image super-resolution,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 5791–5800
2020
Later among the works it cites.
H. Zhang, Q. Yao, M. Yang, Y. Xu, and X. Bai, “Autostr: Efficient backbone search for scene text recognition,” arXiv e-prints , pp. arXiv–2003, 2020
2020
Later among the works it cites.
T. Wang, Y. Zhu, L. Jin, C. Luo, X. Chen, Y. Wu, Q. Wang, and M. Cai, “Decoupled attention network for text recognition,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 34, no. 07, 2020, pp. 12 216–12 224
2020
Later among the works it cites.
S. Fang, H. Xie, Y. Wang, Z. Mao, and Y. Zhang, “Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2021, pp. 7098–7107
2021
Closest in time.
A. K. Bhunia, A. Sain, A. Kumar, S. Ghose, P. N. Chowdhury, and Y.-Z. Song, “Joint visual semantic reasoning: Multi-stage decoder for text recognition,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 14 940–14 949
2021
Closest in time.
A. K. Bhunia, P. N. Chowdhury, A. Sain, and Y.-Z. Song, “Towards the unseen: Iterative text recognition by distilling from errors,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2021, pp. 14 950–14 959
2021
Closest in time.
H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jégou, “Training data-efficient image transformers & distillation through attention,” in International Conference on Machine Learning . PMLR, 2021, pp. 10 347–10 357
2021
Closest in time.
2021
Closest in time.
Z. Jiang, Q. Hou, L. Yuan, Z. Daquan, Y. Shi, X. Jin, A. Wang, and J. Feng, “All tokens matter: Token labeling for training better vision transformers,” in Thirty-Fifth Conference on Neural Information Processing Systems , 2021
2021
Closest in time.
2021
Closest in time.
Z. Dai, B. Cai, Y. Lin, and J. Chen, “Up-detr: Unsupervised pre-training for object detection with transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 1601–1610
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
H. Chen, Y. Wang, T. Guo, C. Xu, Y. Deng, Z. Liu, S. Ma, C. Xu, C. Xu, and W. Gao, “Pre-trained image processing transformer,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2021, pp. 12 299–12 310
2021
Closest in time.
M. Kumar, D. Weissenborn, and N. Kalchbrenner, “Colorization transformer,” in International Conference on Learning Representations , 2021
2021
Closest in time.
R. Yan, L. Peng, S. Xiao, and G. Yao, “Primitive representation learning for scene text recognition,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , June 2021, pp. 284–293
2021
Closest in time.
C. Xue, W. Zhang, Y. Hao, S. Lu, P. Torr, and S. Bai, “Language matters: A weakly supervised vision-language pre-training approach for scene text detection and spotting,” Proceedings of the European Conference on Computer Vision (ECCV) , 2022
2022
Closest in time.
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio, “Show, attend and tell: Neural image caption generation with visual attention,” in International conference on machine learning . PMLR, 2015, pp. 2048–2057
2057
Closest in time.
F. Zhan and S. Lu, “Esir: End-to-end scene text recognition via iterative image rectification,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2019, pp. 2059–2068
2068
Closest in time.