Fetching the paper…
Reading the bibliography…
Most existing video text spotting benchmarks focus on evaluating a single language and scenario with limited data.
The hungarian method for the assignment problem
Harold W Kuhn · 1955
Earlier work this paper cites.
A video text detection and recognition system
Jie Xi, Xian-Sheng Hua, Xiang-Rong Chen, Liu Wenyin, and Hong-Jiang Zhang · 2001
Earlier work this paper cites.
Efficient video text recognition using multiple frame integration
Xian-Sheng Hua, Pei Yin, and Hong-Jiang Zhang · 2002
Earlier work this paper cites.
Text information extraction in images and video: a survey
Keechul Jung, Kwang In Kim, and Anil K Jain · 2004
Earlier work this paper cites.
License plate recognition from still images and video sequences: A survey
Christos-Nikolaos E Anagnostopoulos, Ioannis E Anagnostopoulos, Ioannis D Psoroulas, Vassili Loumos, and Eleftherios Kayafas · 2008
Earlier work this paper cites.
Evaluating multiple object tracking performance: the clear mot metrics
Keni Bernardin and Rainer Stiefelhagen · 2008
Earlier work this paper cites.
Learning to associate: Hybridboosted multi-target tracker for crowded scene
Yuan Li, Chang Huang, and Ram Nevatia · 2009
Earlier work this paper cites.
Using multiple frame integration for the text recognition of video
Jian Yi, Yuxin Peng, and Jianguo Xiao · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Snoopertrack: Text detection and tracking for outdoor videos
Rodrigo Minetto, Nicolas Thome, Matthieu Cord, Neucimar J Leite, and Jorge Stolfi · 2011
Earlier work this paper cites.
Exploiting text-related features for content-based image retrieval
Georg Schroth, Sebastian Hilsenbeck, Robert Huitl, Florian Schweiger, and Eckehard Steinbach · 2011
Earlier work this paper cites.
End-to-end scene text recognition
Kai Wang, Boris Babenko, and Serge Belongie · 2011
Earlier work this paper cites.
Icdar 2013 robust reading competition
Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluis Gomez i Bigorda, Sergi Robles Mestre, Joan Mas, David Fernandez Mota, Jon Almazan Almazan, and Lluis Pere De Las Heras · 2013
Earlier work this paper cites.
Scene text detection via connected component clustering and nontext filtering
Hyung Il Koo and Duck Hoon Kim · 2013
Earlier work this paper cites.
Image retrieval using textual cues
Anand Mishra, Karteek Alahari, and CV Jawahar · 2013
Earlier work this paper cites.
Synthetic data and artificial neural networks for natural scene text recognition
Max Jaderberg, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
Road-sign text recognition architecture for intelligent transportation systems
Abdelhamid Mammeri, El-Hebri Khiari, and Azzedine Boukerche · 2014
Earlier work this paper cites.
Video text detection and recognition: Dataset and benchmark
Phuc Xuan Nguyen, Kai Wang, and Serge Belongie · 2014
Earlier work this paper cites.
Scene text recognition in multiple frames based on text tracking
Xuejian Rong, Chucai Yi, Xiaodong Yang, and Yingli Tian · 2014
Earlier work this paper cites.
Spatial transformer networks
Max Jaderberg, Karen Simonyan, Andrew Zisserman, and Koray Kavukcuoglu · 2015
Earlier work this paper cites.
Icdar 2015 competition on robust reading
Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman Ghosh, Andrew Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ramaseshan Chandrasekhar, Shijian Lu, et al · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov · 2015
Earlier work this paper cites.
A new technique for multi-oriented scene text line detection and tracking in video
Liang Wu, Palaiahnakote Shivakumara, Tong Lu, and Chew Lim Tan · 2015
Earlier work this paper cites.
Icdar 2015 text reading in the wild competition
Xinyu Zhou, Shuchang Zhou, Cong Yao, Zhimin Cao, and Qi Yin · 2015
Earlier work this paper cites.
Synthetic data for text localisation in natural images
Ankush Gupta, Andrea Vedaldi, and Andrew Zisserman · 2016
Earlier work this paper cites.
Tgif: A new dataset and benchmark on animated gif description
Yuncheng Li, Yale Song, Liangliang Cao, Joel Tetreault, Larry Goldberg, Alejandro Jaimes, and Jiebo Luo · 2016
Cited alongside, same era.
Mser-based text detection and communication algorithm for autonomous vehicles
Abdelhamid Mammeri, Azzedine Boukerche, et al · 2016
Cited alongside, same era.
Performance measures and a data set for multi-target, multi-camera tracking
Ergys Ristani, Francesco Solera, Roger Zou, Rita Cucchiara, and Carlo Tomasi · 2016
Cited alongside, same era.
An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition
Baoguang Shi, Xiang Bai, and Cong Yao · 2016
Cited alongside, same era.
Robust scene text recognition with automatic rectification
Baoguang Shi, Xinggang Wang, Pengyuan Lyu, Cong Yao, and Xiang Bai · 2016
Cited alongside, same era.
Generalized intersection over union: A metric and a loss for bounding box regression
Hamid Rezatofighi, Nathan Tsoi, JunYoung Gwak, Amir Sadeghian, Ian Reid, and Silvio Savarese · 2019
Later among the works it cites.
Towards vqa models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach · 2019
Later among the works it cites.
Shape robust text detection with progressive scale expansion network
Wenhai Wang, Enze Xie, Xiang Li, Wenbo Hou, Tong Lu, Gang Yu, and Shuai Shao · 2019
Later among the works it cites.
Vatex: A large-scale, high-quality multilingual dataset for video-and-language research
Xin Wang, Jiawei Wu, Junkun Chen, Lei Li, Yuan-Fang Wang, and William Yang Wang · 2019
Later among the works it cites.
End-to-end object detection with transformers
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andreas Veit, Tomas Matera, Lukas Neumann, Jiri Matas, and Serge Belongie · 2016
Cited alongside, same era.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui · 2016
Cited alongside, same era.
Text detection, tracking and recognition in video: a comprehensive survey
Xu-Cheng Yin, Ze-Yu Zuo, Shu Tian, and Cheng-Lin Liu · 2016
Cited alongside, same era.
Text detection in arabic news video based on SWT operator and convolutional auto-encoders
Oussama Zayene, Mathias Seuret, Sameh Masmoudi Touj, Jean Hennebert, Rolf Ingold, and Najoua Essoukri Ben Amara · 2016
Cited alongside, same era.
Total-text: A comprehensive dataset for scene text detection and recognition
Chee Kheng Ch’ng and Chee Seng Chan · 2017
Cited alongside, same era.
Towards end-to-end text spotting with convolutional recurrent neural networks
Hui Li, Peng Wang, and Chunhua Shen · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Zhanzhan Cheng, Jing Lu, Baorui Zou, Liang Qiao, Yunlu Xu, Shiliang Pu, Yi Niu, Fei Wu, and Shuigeng Zhou · 2020
Later among the works it cites.
An image is worth 16x16 words: Transformers for image recognition at scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al · 2020
Later among the works it cites.
Tvr: A large-scale dataset for video-subtitle moment retrieval
Jie Lei, Licheng Yu, Tamara L Berg, and Mohit Bansal · 2020
Later among the works it cites.
Real-time scene text detection with differentiable binarization
Minghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen, and Xiang Bai · 2020
Later among the works it cites.
Violin: A large-scale dataset for video-and-language inference
Jingzhou Liu, Wenhu Chen, Yu Cheng, Zhe Gan, Licheng Yu, Yiming Yang, and Jingjing Liu · 2020
Later among the works it cites.
Roadtext-1k: Text detection & recognition dataset for driving videos
Sangeeth Reddy, Minesh Mathew, Lluis Gomez, Marçal Rusinol, Dimosthenis Karatzas, and CV Jawahar · 2020
Later among the works it cites.
Textcaps: a dataset for image captioning with reading comprehension
Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach, and Amanpreet Singh · 2020
Later among the works it cites.
Transtrack: Multiple-object tracking with transformer
Peize Sun, Yi Jiang, Rufeng Zhang, Enze Xie, Jinkun Cao, Xinting Hu, Tao Kong, Zehuan Yuan, Changhu Wang, and Ping Luo · 2020
Later among the works it cites.
Hengshuang Zhao, Li Jiang, Jiaya Jia, Philip Torr, and Vladlen Koltun · 2020
Later among the works it cites.
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2020
Later among the works it cites.
Dual encoding for video retrieval by text
Jianfeng Dong, Xirong Li, Chaoxi Xu, Xun Yang, Gang Yang, Xun Wang, and Meng Wang · 2021
Closest in time.
Semantic-aware video text detection
Wei Feng, Fei Yin, Xu-Yao Zhang, and Cheng-Lin Liu · 2021
Closest in time.
Scene text detection and recognition: The deep learning era
Shangbang Long, Xin He, and Cong Yao · 2021
Closest in time.
Master: Multi-aspect non-local network for scene text recognition
Ning Lu, Wenwen Yu, Xianbiao Qi, Yihao Chen, Ping Gong, Rong Xiao, and Xiang Bai · 2021
Closest in time.
Scene text retrieval via joint text detection and similarity learning
Hao Wang, Xiang Bai, Mingkun Yang, Shenggao Zhu, Jing Wang, and Wenyu Liu · 2021
Closest in time.
Transformer meets tracker: Exploiting temporal context for robust visual tracking
Ning Wang, Wengang Zhou, Jie Wang, and Houqiang Li · 2021
Closest in time.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao · 2021
Closest in time.
End-to-end video instance segmentation with transformers
Yuqing Wang, Zhaoliang Xu, Xinlong Wang, Chunhua Shen, Baoshan Cheng, Hao Shen, and Huaxia Xia · 2021
Closest in time.
End-to-end video text detection with online tracking
Hongyuan Yu, Yan Huang, Lihong Pi, Chengquan Zhang, Xuan Li, and Liang Wang · 2021
Closest in time.
Motr: End-to-end multiple-object tracking with transformer
Fangao Zeng, Bin Dong, Tiancai Wang, Cheng Chen, Xiangyu Zhang, and Yichen Wei · 2021
Closest in time.
Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers
Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, Yanwei Fu, Jianfeng Feng, Tao Xiang, Philip HS Torr, et al · 2021
Closest in time.