Fetching the paper…
Reading the bibliography…
Recent video text spotting methods usually require the three-staged pipeline, i.e., detecting text in individual images, recognizing localized text, tracking text streams with post-processing to generate final results.
The hungarian method for the assignment problem
Harold W. Kuhn · 1955
Earlier work this paper cites.
Improvement of video text recognition by character selection
Takeshi Mita and Osamu Hori · 2001
Earlier work this paper cites.
Detecting text in natural scenes with stroke width transform
Boris Epshtein, Eyal Ofek, and Yonatan Wexler · 2010
Earlier work this paper cites.
Text from corners: a novel approach to detect text and caption in videos
Xu Zhao, Kai-Hsiang Lin, Yun Fu, Yuxiao Hu, Yuncai Liu, and Thomas S Huang · 2010
Earlier work this paper cites.
Snoopertrack: Text detection and tracking for outdoor videos
Rodrigo Minetto, Nicolas Thome, Matthieu Cord, Neucimar J Leite, and Jorge Stolfi · 2011
Earlier work this paper cites.
Icdar 2013 robust reading competition
Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluis Gomez i Bigorda, and Sergi Robles Mestre · 2013
Earlier work this paper cites.
Scene text detection via connected component clustering and nontext filtering
Hyung Il Koo and Duck Hoon Kim · 2013
Earlier work this paper cites.
Robust text detection in natural scene images
Xu-Cheng Yin, Xuwang Yin, Kaizhu Huang, and Hong-Wei Hao · 2013
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
Video text detection and recognition: Dataset and benchmark
Phuc Xuan Nguyen, Kai Wang, and Serge Belongie · 2014
Earlier work this paper cites.
Scene text recognition in multiple frames based on text tracking
Xuejian Rong, Chucai Yi, Xiaodong Yang, and Yingli Tian · 2014
Earlier work this paper cites.
Icdar 2015 competition on robust reading
Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman Ghosh, Andrew Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ramaseshan Chandrasekhar, and Shijian Lu · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Unsupervised learning of video representations using lstms
Nitish Srivastava, Elman Mansimov, and Ruslan Salakhudinov · 2015
Earlier work this paper cites.
A new technique for multi-oriented scene text line detection and tracking in video
Liang Wu, Palaiahnakote Shivakumara, Tong Lu, and Chew Lim Tan · 2015
Earlier work this paper cites.
Multi-strategy tracking based text detection in scene videos
Ze-Yu Zuo, Shu Tian, Wei-yi Pei, and Xu-Cheng Yin · 2015
Earlier work this paper cites.
Coco-text: Dataset and benchmark for text detection and recognition in natural images
Andreas Veit, Tomas Matera, Lukas Neumann, Jiri Matas, and Serge Belongie · 2016
Earlier work this paper cites.
Text detection, tracking and recognition in video: a comprehensive survey
Xu-Cheng Yin, Ze-Yu Zuo, Shu Tian, and Cheng-Lin Liu · 2016
Earlier work this paper cites.
Arbitrarily-oriented multi-lingual text detection in video
Vijeta Khare, Palaiahnakote Shivakumara, Raveendran Paramesran, and Michael Blumenstein · 2017
Earlier work this paper cites.
Towards end-to-end text spotting with convolutional recurrent neural networks
Hui Li, Peng Wang, and Chunhua Shen · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
Fractals based multi-oriented text detection system for recognition in mobile video images
Palaiahnakote Shivakumara, Liang Wu, Tong Lu, Chew Lim Tan, Michael Blumenstein, and Basavaraj S Anami · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
End-to-end scene text recognition in videos based on multi frame tracking
Xiaobing Wang, Yingying Jiang, Shuli Yang, Xiangyu Zhu, Wei Li, Pei Fu, Hua Wang, and Zhenbo Luo · 2017
Cited alongside, same era.
East: an efficient and accurate scene text detector
Xinyu Zhou, Cong Yao, He Wen, Yuzhi Wang, Shuchang Zhou, Weiran He, and Jiajun Liang · 2017
Cited alongside, same era.
Rrpn++: Guidance towards more accurate scene text detection
Jianqi Ma · 2020
Later among the works it cites.
Roadtext-1k: Text detection & recognition dataset for driving videos
Sangeeth Reddy, Minesh Mathew, Lluis Gomez, Marçal Rusinol, Dimosthenis Karatzas, and CV Jawahar · 2020
Later among the works it cites.
Transtrack: Multiple-object tracking with transformer
Peize Sun, Yi Jiang, Rufeng Zhang, Enze Xie, Jinkun Cao, Xinting Hu, Tao Kong, Zehuan Yuan, Changhu Wang, and Ping Luo · 2020
Later among the works it cites.
Synthetic-to-real unsupervised domain adaptation for scene text detection in the wild
Weijia Wu, Ning Lu, Enze Xie, Yuxing Wang, Wenwen Yu, Cheng Yang, and Hong Zhou · 2020
Later among the works it cites.
Tracking objects as points
Xingyi Zhou, Vladlen Koltun, and Philipp Krähenbühl · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Fots: Fast oriented text spotting with a unified network
Xuebo Liu, Ding Liang, Shi Yan, Dagui Chen, Yu Qiao, and Junjie Yan · 2018
Cited alongside, same era.
Mask textspotter: An end-to-end trainable neural network for spotting text with arbitrary shapes
Pengyuan Lyu, Minghui Liao, Cong Yao, Wenhao Wu, and Xiang Bai · 2018
Cited alongside, same era.
Mask textspotter: An end-to-end trainable neural network for spotting text with arbitrary shapes
Pengyuan Lyu, Minghui Liao, Cong Yao, Wenhao Wu, and Xiang Bai · 2018
Cited alongside, same era.
Arbitrary-oriented scene text detection via rotation proposals
Jianqi Ma, Weiyuan Shao, Hao Ye, Li Wang, Hong Wang, Yingbin Zheng, and Xiangyang Xue · 2018
Cited alongside, same era.
Scene video text tracking with graph matching
Wei-Yi Pei, Chun Yang, Li-Yu Meng, Jie-Bo Hou, Shu Tian, and Xu-Cheng Yin · 2018
Cited alongside, same era.
Scene text detection and tracking in video with background cues
Lan Wang, Yang Wang, Susu Shan, and Feng Su · 2018
Cited alongside, same era.
Character region awareness for text detection
Youngmin Baek, Bado Lee, Dongyoon Han, Sangdoo Yun, and Hwalsuk Lee · 2019
Cited alongside, same era.
Deformable detr: Deformable transformers for end-to-end object detection
Xizhou Zhu, Weijie Su, Lewei Lu, Bin Li, Xiaogang Wang, and Jifeng Dai · 2020
Later among the works it cites.
Dual encoding for video retrieval by text
Jianfeng Dong, Xirong Li, Chaoxi Xu, Xun Yang, Gang Yang, Xun Wang, and Meng Wang · 2021
Later among the works it cites.
Semantic-aware video text detection
Wei Feng, Fei Yin, Xu-Yao Zhang, and Cheng-Lin Liu · 2021
Later among the works it cites.
Swin transformer: Hierarchical vision transformer using shifted windows
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo · 2021
Later among the works it cites.
Transformer meets tracker: Exploiting temporal context for robust visual tracking
Ning Wang, Wengang Zhou, Jie Wang, and Houqiang Li · 2021
Later among the works it cites.
Towards end-to-end text spotting in natural scenes
Peng Wang, Hui Li, and Chunhua Shen · 2021
Later among the works it cites.
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions
Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, and Ling Shao · 2021
Later among the works it cites.
Pan++: Towards efficient and accurate end-to-end spotting of arbitrarily-shaped text
Wenhai Wang, Enze Xie, Xiang Li, Xuebo Liu, Ding Liang, Yang Zhibo, Tong Lu, and Chunhua Shen · 2021
Later among the works it cites.
A bilingual, openworld video text dataset and end-to-end video text spotter with transformer
Weijia Wu, Debing Zhang, Yuanqiang Cai, Sibo Wang, Jiahong Li, Zhuang Li, Yejun Tang, and Hong Zhou · 2021
Later among the works it cites.
End-to-end video text detection with online tracking
Hongyuan Yu, Yan Huang, Lihong Pi, Chengquan Zhang, Xuan Li, and Liang Wang · 2021
Later among the works it cites.
Motr: End-to-end multiple-object tracking with transformer
Fangao Zeng, Bin Dong, Tiancai Wang, Cheng Chen, Xiangyu Zhang, and Yichen Wei · 2021
Later among the works it cites.
Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers
Sixiao Zheng, Jiachen Lu, Hengshuang Zhao, Xiatian Zhu, Zekun Luo, Yabiao Wang, and Fu · 2021
Later among the works it cites.
Neighborhood-adaptive structure augmented metric learning
Pandeng Li, Yan Li, Hongtao Xie, and Lei Zhang · 2022
Closest in time.
Tracking objects as pixel-wise distributions, 2022
Zelin Zhao, Ze Wu, Yueqing Zhuang, Boxun Li, and Jiaya Jia · 2022
Closest in time.