Fetching the paper…
Reading the bibliography…
Existing text recognition methods usually need large-scale training data.
The IAM-database: an English sentence database for offline handwriting recognition
Urs-Viktor Marti and Horst Bunke. 2002 · 2002
Earlier work this paper cites.
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C. Bovik, Hamid R. Sheikh, and Eero P. Simoncelli. 2004 · 2004
Earlier work this paper cites.
Histograms of Oriented Gradients for Human Detection. In Proc. Conf. Comput. Vision Pattern Recognition . IEEE Computer Society, 886–893
Navneet Dalal and Bill Triggs. 2005 · 2005
Earlier work this paper cites.
Dimensionality Reduction by Learning an Invariant Mapping. In Proc. Conf. Comput. Vision Pattern Recognition . IEEE Computer Society, 1735–1742
Raia Hadsell, Sumit Chopra, and Yann LeCun. 2006 · 2006
Earlier work this paper cites.
End-to-end scene text recognition. In Proc. Int. Conf. Comput. Vision . IEEE Computer Society, 1457–1464
Kai Wang, Boris Babenko, and Serge Belongie. 2011 · 2011
Earlier work this paper cites.
Top-down and bottom-up cues for scene text recognition. In Proc. Conf. Comput. Vision Pattern Recognition . IEEE Computer Society
Anand Mishra, Karteek Alahari, and C. V. Jawahar. 2012 · 2012
Earlier work this paper cites.
ICDAR 2013 robust reading competition. In Proc. ICDAR . 1484–1493
Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluis Gomez i Bigorda, Sergi Robles Mestre, Joan Mas, David Fernandez Mota, Jon Almazan Almazan, and Lluis Pere de las Heras. 2013 · 2013
Earlier work this paper cites.
CVL-DataBase: An Off-Line Database for Writer Retrieval, Writer Identification and Word Spotting. In Proc. ICDAR . IEEE Computer Society, 560–564
Florian Kleber, Stefan Fiel, Markus Diem, and Robert Sablatnig. 2013 · 2013
Earlier work this paper cites.
Recognizing text with perspective distortion in natural scenes. In Proc. Int. Conf. Comput. Vision . IEEE Computer Society, 569–576
Trung Quy Phan, Palaiahnakote Shivakumara, Shangxuan Tian, and Chew Lim Tan. 2013 · 2013
Earlier work this paper cites.
Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition
Max Jaderberg, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
A robust arbitrary text detection system for natural scene images
Anhar Risnumawan, Palaiahnakote Shivakumara, Chee Seng Chan, and Chew Lim Tan. 2014 · 2014
Earlier work this paper cites.
Unsupervised Visual Representation Learning by Context Prediction. In Proc. Int. Conf. Comput. Vision . IEEE Computer Society, 1422–1430
Carl Doersch, Abhinav Gupta, and Alexei A. Efros. 2015 · 2015
Earlier work this paper cites.
ICDAR 2015 competition on Robust Reading. In Proc. ICDAR . 1156–1160
Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman K. Ghosh, Andrew D. Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ramaseshan Chandrasekhar, Shijian Lu, Faisal Shafait, Seiichi Uchida, and Ernest Valveny. 2015 · 2015
Earlier work this paper cites.
Image Super-Resolution Using Deep Convolutional Networks
Chao Dong, Chen Change Loy, Kaiming He, and Xiaoou Tang. 2016 · 2016
Earlier work this paper cites.
Discriminative Unsupervised Feature Learning with Exemplar Convolutional Neural Networks
Alexey Dosovitskiy, Philipp Fischer, Jost Tobias Springenberg, Martin A. Riedmiller, and Thomas Brox. 2016 · 2016
Earlier work this paper cites.
Synthetic Data for Text Localisation in Natural Images. In Proc. Conf. Comput. Vision Pattern Recognition . IEEE Computer Society, 2315–2324
Ankush Gupta, Andrea Vedaldi, and Andrew Zisserman. 2016 · 2016
Earlier work this paper cites.
Reading Scene Text in Deep Convolutional Sequences. In Proc. AAAI
Pan He, Weilin Huang, Yu Qiao, Chen Change Loy, and Xiaoou Tang. 2016 · 2016
Earlier work this paper cites.
Reading Text in the Wild with Convolutional Neural Networks
Max Jaderberg, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2016 · 2016
Earlier work this paper cites.
COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images
Andreas Veit, Tomas Matera, Lukas Neumann, Jiri Matas, and Serge J. Belongie. 2016 · 2016
Earlier work this paper cites.
Total-Text: A Comprehensive Dataset for Scene Text Detection and Recognition. In Proc. ICDAR . 935–942
Chee Kheng Chng and Chee Seng Chan. 2017 · 2017
Earlier work this paper cites.
Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network
Christian Ledig, Lucas Theis, Ferenc Huszár, Jose Caballero, Andrew P. Aitken, Alykhan Tejani, Johannes Totz, Zehan Wang, and Wenzhe Shi. 2017 · 2017
Earlier work this paper cites.
Pointer Sentinel Mixture Models. In Int. Conf. on Learning Representations
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017 · 2017
Cited alongside, same era.
An End-to-End Trainable Neural Network for Image-Based Sequence Recognition and Its Application to Scene Text Recognition
Baoguang Shi, Xiang Bai, and Cong Yao. 2017 · 2017
Cited alongside, same era.
Accurate recognition of words in scenes without character segmentation using recurrent neural network
Bolan Su and Shijian Lu. 2017 · 2017
Cited alongside, same era.
Attention is All you Need. In Neural Inform. Process. Syst. 5998–6008
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Attention-based extraction of structured information from street view imagery. In Proc. ICDAR , Vol. 1. IEEE, 844–850
Zbigniew Wojna, Alexander N Gorban, Dar-Shyang Lee, Kevin Murphy, Qian Yu, Yeqing Li, and Julian Ibarz. 2017 · 2017
Cited alongside, same era.
RobustScanner: Dynamically Enhancing Positional Clues for Robust Text Recognition. In European Conf. Comput. Vision . 135–151
Xiaoyu Yue, Zhanghui Kuang, Chenhao Lin, Hongbin Sun, and Wayne Zhang. 2020 · 2020
Later among the works it cites.
Hui Zhang, Quanming Yao, Mingkun Yang, Yongchao Xu, and Xiang Bai. 2020 · 2020
Later among the works it cites.
Sequence-to-Sequence Contrastive Learning for Text Recognition. In Proc. Conf. Comput. Vision Pattern Recognition . IEEE Computer Society, 15302–15312
Aviad Aberdam, Ron Litman, Shahar Tsiper, Oron Anschel, Ron Slossberg, Shai Mazor, R. Manmatha, and Pietro Perona. 2021 · 2021
Later among the works it cites.
Joint Visual Semantic Reasoning: Multi-Stage Decoder for Text Recognition. In Proc. Int. Conf. Comput. Vision . 14920–14929
Ayan Kumar Bhunia, Aneeshan Sain, Amandeep Kumar, Shuvozit Ghose, Pinaki Nath Chowdhury, and Yi-Zhe Song. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Learning to read irregular text with attention mechanisms. In Proc. IJCAI . 3280–3286
Xiao Yang, Dafang He, Zihan Zhou, Daniel Kifer, and C Lee Giles. 2017 · 2017
Cited alongside, same era.
TextBoxes++: A Single-Shot Oriented Scene Text Detector
Minghui Liao, Baoguang Shi, and Xiang Bai. 2018 · 2018
Cited alongside, same era.
Representation Learning with Contrastive Predictive Coding
Aäron van den Oord, Yazhe Li, and Oriol Vinyals. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2019, Minneapolis, MN, USA, June 2-7, 2019, Volume 1 (Long and Short Papers) , Jill Burstein, Christy Doran, and Thamar Solorio (Eds.). 4171–4186
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Mask TextSpotter: An End-to-End Trainable Neural Network for Spotting Text with Arbitrary Shapes
Minghui Liao, Pengyuan Lyu, Minghang He, Cong Yao, Wenhao Wu, and Xiang Bai. 2021 · 2019
Cited alongside, same era.
Scene Text Recognition from Two-Dimensional Perspective. In AAAI Conf. on Artificial Intelligence . AAAI Press, 8714–8721
Minghui Liao, Jian Zhang, Zhaoyi Wan, Fengming Xie, Jiajun Liang, Pengyuan Lyu, Cong Yao, and Xiang Bai. 2019 · 2019
Cited alongside, same era.
Curved scene text detection via transverse and longitudinal sequence connection
Yuliang Liu, Lianwen Jin, Shuaitao Zhang, Canjie Luo, and Sheng Zhang. 2019 · 2019
Cited alongside, same era.
Scene Text Telescope: Text-Focused Scene Image Super-Resolution
Jingye Chen, Bin Li, and X. Xue. 2021a · 2021
Later among the works it cites.
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. In Int. Conf. on Learning Representations . OpenReview.net
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021 · 2021
Later among the works it cites.
Read Like Humans: Autonomous, Bidirectional and Iterative Language Modeling for Scene Text Recognition. In Proc. Conf. Comput. Vision Pattern Recognition . Computer Vision Foundation / IEEE, 7098–7107
Shancheng Fang, Hongtao Xie, Yuxin Wang, Zhendong Mao, and Yongdong Zhang. 2021 · 2021
Later among the works it cites.
Masked Autoencoders Are Scalable Vision Learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross B. Girshick. 2021 · 2021
Later among the works it cites.
MASTER: Multi-aspect non-local network for scene text recognition
Ning Lu, Wenwen Yu, Xianbiao Qi, Yihao Chen, Ping Gong, Rong Xiao, and Xiang Bai. 2021 · 2021
Later among the works it cites.
TextOCR: Towards Large-Scale End-to-End Reasoning for Arbitrary-Shaped Scene Text. In Proc. Conf. Comput. Vision Pattern Recognition . IEEE Computer Society, 8802–8812
Amanpreet Singh, Guan Pang, Mandy Toh, Jing Huang, Wojciech Galuba, and Tal Hassner. 2021 · 2021
Later among the works it cites.
From Two to One: A New Scene Text Recognizer with Visual Language Modeling Network. In Proc. Int. Conf. Comput. Vision . IEEE Computer Society, 14174–14183
Yuxin Wang, Hongtao Xie, Shancheng Fang, Jing Wang, Shenggao Zhu, and Yongdong Zhang. 2021 · 2021
Later among the works it cites.
Masked Feature Prediction for Self-Supervised Visual Pre-Training
Chen Wei, Haoqi Fan, Saining Xie, Chao-Yuan Wu, Alan L. Yuille, and Christoph Feichtenhofer. 2021 · 2021
Later among the works it cites.
SimMIM: A Simple Framework for Masked Image Modeling
Zhenda Xie, Zheng Zhang, Yue Cao, Yutong Lin, Jianmin Bao, Zhuliang Yao, Qi Dai, and Han Hu. 2021 · 2021
Later among the works it cites.
Rethinking Text Segmentation: A Novel Dataset and a Text-Specific Refinement Approach. In Proc. Conf. Comput. Vision Pattern Recognition . IEEE Computer Society, 12045–12055
Xingqian Xu, Zhifei Zhang, Zhaowen Wang, Brian Price, Zhonghao Wang, and Humphrey Shi. 2021 · 2021
Later among the works it cites.
Primitive Representation Learning for Scene Text Recognition. In Proc. Conf. Comput. Vision Pattern Recognition . 284–293
Ruijie Yan, Liangrui Peng, Shanyu Xiao, and Gang Yao. 2021 · 2021
Later among the works it cites.
TAP: Text-Aware Pre-training for Text-VQA and Text-Caption. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 8751–8761
Zhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin, Dinei Florencio, Lijuan Wang, Cha Zhang, Lei Zhang, and Jiebo Luo. 2021 · 2021
Later among the works it cites.
SPIN: Structure-Preserving Inner Offset Network for Scene Text Recognition. In AAAI Conf. on Artificial Intelligence . 3305–3314
Chengwei Zhang, Yunlu Xu, Zhanzhan Cheng, Shiliang Pu, Yi Niu, Fei Wu, and Futai Zou. 2021 · 2021
Later among the works it cites.
Perceiving Stroke-Semantic Context: Hierarchical Contrastive Learning for Robust Scene Text Recognition. In AAAI Conf. on Artificial Intelligence
Hao Liu, Bin Wang, Zhimin Bao, Mobai Xue, Sheng Kang, Deqiang Jiang, Yinsong Liu, and Bo Ren. 2022 · 2022
Closest in time.
Three things everyone should know about Vision Transformers
Hugo Touvron, Matthieu Cord, Alaaeldin El-Nouby, Jakob Verbeek, and Herv’e J’egou. 2022 · 2022
Closest in time.
ASTER: An Attentional Scene Text Recognizer with Flexible Rectification
Baoguang Shi, Mingkun Yang, Xinggang Wang, Pengyuan Lyu, Cong Yao, and Xiang Bai. 2019 · 2048
Closest in time.