Fetching the paper…
Reading the bibliography…
The Transformer-based encoder-decoder framework is becoming popular in scene text recognition, largely because it naturally integrates recognition clues from both visual and semantic domains.
Wang K, Babenko B, Belongie S (2011) End-to-end scene text recognition. In: ICCV, pp 1457–1464
2011
Earlier work this paper cites.
Mishra A, Alahari K, Jawahar C (2012) Scene text recognition using higher order language priors. In: BMVC, pp 1–11
2012
Earlier work this paper cites.
Phan TQ, Shivakumara P, Tian S, Tan CL (2013) Recognizing text with perspective distortion in natural scenes. In: ICCV, pp 569–576
2013
Earlier work this paper cites.
Bai J, Chen Z, Feng B, Xu B (2014) Chinese image text recognition on grayscale pixels. In: 2014 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, pp 1380–1384
2014
Earlier work this paper cites.
Jaderberg M, Simonyan K, Vedaldi A, Zisserman A (2014) Synthetic data and artificial neural networks for natural scene text recognition. arXiv preprint arXiv:14062227
2014
Earlier work this paper cites.
Risnumawan A, Shivakumara P, Chan CS, Tan CL (2014) A robust arbitrary text detection system for natural scene images. ESA 41(18):8027–8048
2014
Earlier work this paper cites.
Ye Q, Doermann D (2014) Text detection and recognition in imagery: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 37(7):1480–1500
2014
Earlier work this paper cites.
Karatzas D, Gomez-Bigorda L, Nicolaou A, Ghosh S, Bagdanov A, Iwamura M, Matas J, Neumann L, Chandrasekhar VR, Lu S, et al. (2015) Icdar 2015 competition on robust reading. In: ICDAR, pp 1156–1160
2015
Earlier work this paper cites.
Rodriguez-Serrano JA, Gordo A, Perronnin F (2015) Label embedding: A frugal baseline for text recognition. International Journal of Computer Vision 113(3):193–207
2015
Earlier work this paper cites.
Gupta A, Vedaldi A, Zisserman A (2016) Synthetic data for text localisation in natural images. In: CVPR, pp 2315–2324
2016
Earlier work this paper cites.
He P, Huang W, Qiao Y, Loy CC, Tang X (2016) Reading scene text in deep convolutional sequences. In: AAAI, pp 3501–3508
2016
Earlier work this paper cites.
Jaderberg M, Simonyan K, Vedaldi A, Zisserman A (2016) Reading text in the wild with convolutional neural networks. International Journal of Computer Vision 116(1):1–20
2016
Earlier work this paper cites.
Lee CY, Osindero S (2016) Recursive recurrent nets with attention modeling for ocr in the wild. In: CVPR, pp 2231–2239
2016
Earlier work this paper cites.
Cheng Z, Bai F, Xu Y, Zheng G, Pu S, Zhou S (2017) Focusing attention: Towards accurate text recognition in natural images. In: ICCV, pp 5076–5084
2017
Earlier work this paper cites.
Li Y, Qi H, Dai J, Ji X, Wei Y (2017) Fully convolutional instance-aware semantic segmentation. In: CVPR, pp 2359–2367
2017
Earlier work this paper cites.
Nayef N, Yin F, Bizid I, Choi H, Feng Y, Karatzas D, Luo Z, Pal U, Rigaud C, Chazalon J, et al. (2017) Icdar2017 robust reading challenge on multi-lingual scene text detection and script identification-rrc-mlt. In: ICDAR, IEEE, vol 1, pp 1454–1459
2017
Earlier work this paper cites.
Shi B, Bai X, Yao C (2017) An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence 39(11):2298–2304
2017
Earlier work this paper cites.
Su B, Lu S (2017) Accurate recognition of words in scenes without character segmentation using recurrent neural network. Pattern Recognition 63:397–405
2017
Earlier work this paper cites.
Vaswani A, Shazeer N, Parmar N, Uszkoreit J, Jones L, Gomez AN, Kaiser L, Polosukhin I (2017) Attention is all you need. In: NIPS, pp 5998–6008
2017
Earlier work this paper cites.
Zhang Y, Gueguen L, Zharkov I, Zhang P, Seifert K, Kadlec B (2017) Uber-text: A large-scale dataset for optical character recognition from street-level imagery. In: SUNw: Scene Understanding Workshop-CVPR
2017
Earlier work this paper cites.
Cheng Z, Xu Y, Bai F, Niu Y, Pu S, Zhou S (2018) Aon: Towards arbitrarily-oriented text recognition. In: CVPR, pp 5571–5579
2018
Cited alongside, same era.
Lyu P, Liao M, Yao C, Wu W, Bai X (2018) Mask textspotter: An end-to-end trainable neural network for spotting text with arbitrary shapes. In: ECCV, pp 67–83
2018
Cited alongside, same era.
Baek J, Kim G, Lee J, Park S, Han D, Yun S, Oh SJ, Lee H (2019) What is wrong with scene text recognition model comparisons? dataset and model analysis. In: ICCV, pp 4714–4722
2019
Cited alongside, same era.
Chng CK, Liu Y, Sun Y, Ng CC, Luo C, Ni Z, Fang C, Zhang S, Han J, Ding E, et al. (2019) Icdar2019 robust reading challenge on arbitrary-shaped text-rrc-art. In: ICDAR, IEEE, pp 1571–1576
2019
Cited alongside, same era.
Devlin J, Chang MW, Lee K, Toutanova K (2019) Bert: Pre-training of deep bidirectional transformers for language understanding. In: NAACL-HLT
Yu D, Li X, Zhang C, Liu T, Han J, Liu J, Ding E (2020) Towards accurate scene text recognition with semantic reasoning networks. In: CVPR, pp 12113–12122
2020
Later among the works it cites.
Yue X, Kuang Z, Lin C, Sun H, Zhang W (2020) Robustscanner: Dynamically enhancing positional clues for robust text recognition. In: ECCV, pp 135–151
2020
Later among the works it cites.
Baek J, Matsui Y, Aizawa K (2021) What if we only use real datasets for scene text recognition? toward scene text recognition with fewer labels. In: CVPR, pp 3113–3122
2021
Closest in time.
Bhunia AK, Sain A, Kumar A, Ghose S, Chowdhury PN, Song YZ (2021) Joint visual semantic reasoning: Multi-stage decoder for text recognition. In: ICCV, pp 14920–14929
2021
Closest in time.
Chen X, Jin L, Zhu Y, Luo C, Wang T (2021) Text recognition in the wild: A survey. ACM Computing Surveys (CSUR) 54(2):1–35
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2019
Cited alongside, same era.
Lan Z, Chen M, Goodman S, Gimpel K, Sharma P, Soricut R (2019) Albert: A lite bert for self-supervised learning of language representations. arXiv preprint arXiv:190911942
2019
Cited alongside, same era.
Li H, Wang P, Shen C, Zhang G (2019) Show, attend and read: A simple and strong baseline for irregular text recognition. In: AAAI, vol 33, pp 8610–8617
2019
Cited alongside, same era.
Liao M, Lyu P, He M, Yao C, Wu W, Bai X (2019) Mask textspotter: An end-to-end trainable neural network for spotting text with arbitrary shapes. IEEE Transactions on Pattern Analysis and Machine Intelligence 43(2):532–548
2019
Cited alongside, same era.
Liao M, Zhang J, Wan Z, Xie F, Liang J, Lyu P, Yao C, Bai X (2019) Scene text recognition from two-dimensional perspective. In: AAAI, vol 33, pp 8714–8721
2019
Cited alongside, same era.
Luo C, Jin L, Sun Z (2019) MORAN: A multi-object rectified attention network for scene text recognition. Pattern Recognition 90:109–118
2019
Cited alongside, same era.
Lyu P, Yang Z, Leng X, Wu X, Li R, Shen X (2019) 2d attentional irregular scene text recognizer. arXiv preprint arXiv:190605708
2019
Cited alongside, same era.
Nayef N, Patel Y, Busta M, Chowdhury PN, Karatzas D, Khlif W, Matas J, Pal U, Burie JC, Liu Cl, et al. (2019) Icdar2019 robust reading challenge on multi-lingual scene text detection and recognition—rrc-mlt-2019. In: ICDAR, IEEE, pp 1582–1587
2019
Cited alongside, same era.
2021
Closest in time.
Fang S, Xie H, Wang Y, Mao Z, Zhang Y (2021) Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition. In: CVPR, pp 7094–7103
2021
Closest in time.
Long S, He X, Yao C (2021) Scene text detection and recognition: The deep learning era. International Journal of Computer Vision 129(1):161–184
2021
Closest in time.
Luo C, Lin Q, Liu Y, Jin L, Shen C (2021) Separating content from style using adversarial learning for recognizing text in the wild. International Journal of Computer Vision 129(4):960–976
2021
Closest in time.
Nguyen N, Nguyen T, Tran V, Tran MT, Ngo TD, Nguyen TH, Hoai M (2021) Dictionary-guided scene text recognition. In: CVPR, pp 7383–7392
2021
Closest in time.
Wang Y, Xie H, Fang S, Wang J, Zhu S, Zhang Y (2021) From two to one: A new scene text recognizer with visual language modeling network. In: ICCV, pp 14174–14183
2021
Closest in time.
Yan R, Peng L, Xiao S, Yao G (2021) Primitive representation learning for scene text recognition. In: CVPR, pp 284–293
2021
Closest in time.
Bautista D, Atienza R (2022) Scene text recognition with permuted autoregressive sequence models. In: ECCV, Springer, pp 178–196
2022
Closest in time.
Da C, Wang P, Yao C (2022) Levenshtein ocr. In: ECCV, Springer, pp 322–338
2022
Closest in time.
Du Y, Chen Z, Jia C, Yin X, Zheng T, Li C, Du Y, Jiang YG (2022) Svtr: Scene text recognition with a single visual model. In: IJCAI, pp 884–890
2022
Closest in time.
Fang S, Mao Z, Xie H, Wang Y, Yan C, Zhang Y (2022) Abinet++: Autonomous, bidirectional and iterative language modeling for scene text spotting. IEEE Transactions on Pattern Analysis and Machine Intelligence
2022
Closest in time.
He Y, Chen C, Zhang J, Liu J, He F, Wang C, Du B (2022) Visual semantics allow for textual reasoning better in scene text recognition. In: AAAI, pp 888–896
2022
Closest in time.
Peng D, Jin L, Liu Y, Luo C, Lai S (2022) Pagenet: Towards end-to-end weakly supervised page-level handwritten chinese text recognition. International Journal of Computer Vision 130(11):2623–2645
2022
Closest in time.
Xie X, Fu L, Zhang Z, Wang Z, Bai X (2022) Toward understanding wordart: Corner-guided transformer for scene text recognition. In: ECCV, Springer, pp 303–321
2022
Closest in time.
Shi B, Yang M, Wang X, Lyu P, Yao C, Bai X (2018) Aster: An attentional scene text recognizer with flexible rectification. IEEE Transactions on Pattern Analysis and Machine Intelligence 41(9):2035–2048
2048
Closest in time.
Zhan F, Lu S (2019) Esir: End-to-end scene text recognition via iterative image rectification. In: CVPR, pp 2059–2068
2068
Closest in time.