Fetching the paper…
Reading the bibliography…
Text images contain both visual and linguistic information.
Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks
Alex Graves, Santiago Fernández, Faustino J. Gomez, and Jürgen Schmidhuber · 2006
Earlier work this paper cites.
End-to-end scene text recognition
Kai Wang, Boris Babenko, and Serge J. Belongie · 2011
Earlier work this paper cites.
Scene text recognition using higher order language priors
Anand Mishra, Karteek Alahari, and C. V. Jawahar · 2012
Earlier work this paper cites.
ICDAR 2013 robust reading competition
Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluis Gomez i Bigorda, Sergi Robles Mestre, Joan Mas, David Fernández Mota, Jon Almazán, and Lluís-Pere de las Heras · 2013
Earlier work this paper cites.
Recognizing text with perspective distortion in natural scenes
Trung Quy Phan, Palaiahnakote Shivakumara, Shangxuan Tian, and Chew Lim Tan · 2013
Earlier work this paper cites.
Synthetic data and artificial neural networks for natural scene text recognition
Max Jaderberg, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
Deep features for text spotting
Max Jaderberg, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
A robust arbitrary text detection system for natural scene images
Anhar Risnumawan, Palaiahnakote Shivakumara, Chee Seng Chan, and Chew Lim Tan · 2014
Earlier work this paper cites.
Spatial transformer networks
Max Jaderberg, Karen Simonyan, Andrew Zisserman, and Koray Kavukcuoglu · 2015
Earlier work this paper cites.
ICDAR 2015 competition on robust reading
Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman K. Ghosh, Andrew D. Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ramaseshan Chandrasekhar, Shijian Lu, Faisal Shafait, Seiichi Uchida, and Ernest Valveny · 2015
Earlier work this paper cites.
Synthetic data for text localisation in natural images
Ankush Gupta, Andrea Vedaldi, and Andrew Zisserman · 2016
Earlier work this paper cites.
Reading text in the wild with convolutional neural networks
Max Jaderberg and Karen Simonyan and Andrea Vedaldi and Andrew Zisserman · 2016
Earlier work this paper cites.
Accurate, large minibatch SGD: training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross B. Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
SGDR: Stochastic Gradient Descent with Warm Restarts
Ilya Loshchilov and Frank Hutter · 2017
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2017
Earlier work this paper cites.
An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition
Baoguang Shi, Xiang Bai, and Cong Yao · 2017
Earlier work this paper cites.
ICDAR2017 competition on reading chinese text in the wild (RCTW-17)
Baoguang Shi, Cong Yao, Minghui Liao, Mingkun Yang, Pei Xu, Linyan Cui, Serge J. Belongie, Shijian Lu, and Xiang Bai · 2017
Earlier work this paper cites.
ICPR2018 contest on robust reading for multi-type web images
Mengchao He, Yuliang Liu, Zhibo Yang, Sheng Zhang, Canjie Luo, Feiyu Gao, Qi Zheng, Yongpan Wang, Xin Zhang, and Lianwen Jin · 2018
Earlier work this paper cites.
Mask textspotter: An end-to-end trainable neural network for spotting text with arbitrary shapes
Pengyuan Lyu, Minghui Liao, Cong Yao, Wenhao Wu, and Xiang Bai · 2018
Earlier work this paper cites.
ICDAR2019 robust reading challenge on arbitrary-shaped text - rrc-art
Chee Kheng Chng, Errui Ding, Jingtuo Liu, Dimosthenis Karatzas, Chee Seng Chan, Lianwen Jin, Yuliang Liu, Yipeng Sun, Chun Chet Ng, Canjie Luo, Zihan Ni, ChuanMing Fang, Shuaitao Zhang, and Junyu Han · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Show, attend and read: A simple and strong baseline for irregular text recognition
Hui Li, Peng Wang, Chunhua Shen, and Guyu Zhang · 2019
Earlier work this paper cites.
Scene text recognition from two-dimensional perspective
Minghui Liao, Jian Zhang, Zhaoyi Wan, Fengming Xie, Jiajun Liang, Pengyuan Lyu, Cong Yao, and Xiang Bai · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Cited alongside, same era.
2d attentional irregular scene text recognizer
Pengyuan Lyu, Zhicheng Yang, Xinhang Leng, Xiaojun Wu, Ruiyu Li, and Xiaoyong Shen · 2019
Cited alongside, same era.
ASTER: an attentional scene text recognizer with flexible rectification
Baoguang Shi, Mingkun Yang, Xinggang Wang, Pengyuan Lyu, Cong Yao, and Xiang Bai · 2019
Cited alongside, same era.
ICDAR 2019 competition on large-scale street view text with partial labeling - RRC-LSVT
Yipeng Sun, Dimosthenis Karatzas, Chee Seng Chan, Lianwen Jin, Zihan Ni, Chee Kheng Chng, Yuliang Liu, Canjie Luo, Chun Chet Ng, Junyu Han, Errui Ding, and Jingtuo Liu · 2019
Cited alongside, same era.
A large chinese text dataset in the wild
Tai-Ling Yuan, Zhe Zhu, Kun Xu, Cheng-Jun Li, Tai-Jiang Mu, and Shi-Min Hu · 2019
Cited alongside, same era.
Trocr: Transformer-based optical character recognition with pre-trained models
Minghao Li, Tengchao Lv, Lei Cui, Yijuan Lu, Dinei A. F. Florêncio, Cha Zhang, Zhoujun Li, and Furu Wei · 2021
Later among the works it cites.
Pimnet: A parallel, iterative and mimicking network for scene text recognition
Zhi Qiao, Yu Zhou, Jin Wei, Wei Wang, Yuan Zhang, Ning Jiang, Hongbin Wang, and Weiping Wang · 2021
Later among the works it cites.
Learning transferable visual models from natural language supervision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever · 2021
Later among the works it cites.
Textocr: Towards large-scale end-to-end reasoning for arbitrary-shaped scene text
Amanpreet Singh, Guan Pang, Mandy Toh, Jing Huang, Wojciech Galuba, and Tal Hassner · 2021
Later among the works it cites.
UFO: A unified transformer for vision-language representation learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
ICDAR 2019 robust reading challenge on reading chinese text on signboard
Rui Zhang, Mingkun Yang, Xiang Bai, Baoguang Shi, Dimosthenis Karatzas, Shijian Lu, C. V. Jawahar, Yongsheng Zhou, Qianyi Jiang, Qi Song, Nan Li, Kai Zhou, Lei Wang, Dong Wang, and Minghui Liao · 2019
Cited alongside, same era.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Cited alongside, same era.
Improved baselines with momentum contrastive learning
Xinlei Chen, Haoqi Fan, Ross B. Girshick, and Kaiming He · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross B. Girshick · 2020
Cited alongside, same era.
SEED: semantics enhanced encoder-decoder framework for scene text recognition
Zhi Qiao, Yu Zhou, Dongbao Yang, Yucan Zhou, and Weiping Wang · 2020
Cited alongside, same era.
Textscanner: Reading characters in order for robust scene text recognition
Zhaoyi Wan, Minghang He, Haoran Chen, Xiang Bai, and Cong Yao · 2020
Cited alongside, same era.
Layoutlm: Pre-training of text and layout for document image understanding
Yiheng Xu, Minghao Li, Lei Cui, Shaohan Huang, Furu Wei, and Ming Zhou · 2020
Cited alongside, same era.
Jianfeng Wang, Xiaowei Hu, Zhe Gan, Zhengyuan Yang, Xiyang Dai, Zicheng Liu, Yumao Lu, and Lijuan Wang · 2021
Later among the works it cites.
Vlmo: Unified vision-language pre-training with mixture-of-modality-experts
Wenhui Wang, Hangbo Bao, Li Dong, and Furu Wei · 2021
Later among the works it cites.
From two to one: A new scene text recognizer with visual language modeling network
Yuxin Wang, Hongtao Xie, Shancheng Fang, Jing Wang, Shenggao Zhu, and Yongdong Zhang · 2021
Later among the works it cites.
I2C2W: image-to-character-to-word transformers for accurate scene text recognition
Chuhui Xue, Shijian Lu, Song Bai, Wenqing Zhang, and Changhu Wang · 2021
Later among the works it cites.
TAP: text-aware pre-training for text-vqa and text-caption
Zhengyuan Yang, Yijuan Lu, Jianfeng Wang, Xi Yin, Dinei Florêncio, Lijuan Wang, Cha Zhang, Lei Zhang, and Jiebo Luo · 2021
Later among the works it cites.
Beit: BERT pre-training of image transformers
Hangbo Bao, Li Dong, Songhao Piao, and Furu Wei · 2022
Closest in time.
Scene text recognition with permuted autoregressive sequence models
Darwin Bautista and Rowel Atienza · 2022
Closest in time.
Context autoencoder for self-supervised representation learning
Xiaokang Chen, Mingyu Ding, Xiaodi Wang, Ying Xin, Shentong Mo, Yunhao Wang, Shumin Han, Ping Luo, Gang Zeng, and Jingdong Wang · 2022
Closest in time.
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross B. Girshick · 2022
Closest in time.
BLIP: bootstrapping language-image pre-training for unified vision-language understanding and generation
Junnan Li, Dongxu Li, Caiming Xiong, and Steven C. H. Hoi · 2022
Closest in time.
Perceiving stroke-semantic context: Hierarchical contrastive learning for robust scene text recognition
Hao Liu, Bin Wang, Zhimin Bao, Mobai Xue, Sheng Kang, Deqiang Jiang, Yinsong Liu, and Bo Ren · 2022
Closest in time.
Vision-language pre-training for boosting scene text detectors
Sibo Song, Jianqiang Wan, Zhibo Yang, Jun Tang, Wenqing Cheng, Xiang Bai, and Cong Yao · 2022
Closest in time.
Multi-granularity prediction for scene text recognition
Peng Wang, Cheng Da, and Cong Yao · 2022
Closest in time.
Reading and writing: Discriminative and generative modeling for self-supervised text recognition
Mingkun Yang, Minghui Liao, Pu Lu, Jing Wang, Shenggao Zhu, Hualin Luo, Qi Tian, and Xiang Bai · 2022
Closest in time.
Coca: Contrastive captioners are image-text foundation models
Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu · 2022
Closest in time.
Context-based contrastive learning for scene text recognition
Xinyun Zhang, Binwu Zhu, Xufeng Yao, Qi Sun, Ruiyu Li, and Bei Yu · 2022
Closest in time.
Pushing the performance limit of scene text recognizer without human annotation
Caiyuan Zheng, Hui Li, Seon-Min Rhee, Seungju Han, Jae-Joon Han, and Peng Wang · 2022
Closest in time.
Revisiting scene text recognition: A data perspective
Qing Jiang, Jiapeng Wang, Dezhi Peng, Chongyu Liu, and Lianwen Jin · 2023
Closest in time.