Fetching the paper…
Reading the bibliography…
A crucial component for the scene text based reasoning required for TextVQA and TextCaps datasets involve detecting and recognizing text present in the images using an optical character recognition (OCR) system.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei Jing Zhu · 2002
Earlier work this paper cites.
Icdar 2003 robust reading competitions
Simon M Lucas, Alex Panaretos, Luis Sosa, Anthony Tang, Shirley Wong, and Robert Young · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
ImageNet: A Large-Scale Hierarchical Image Database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Vizwiz: nearly real-time answers to visual questions
Jeffrey P Bigham, Chandrika Jayant, Hanjie Ji, Greg Little, Andrew Miller, Robert C Miller, Robin Miller, Aubrey Tatarowicz, Brandyn White, Samual White, et al · 2010
Earlier work this paper cites.
End-to-end scene text recognition
Kai Wang, Boris Babenko, and Serge Belongie · 2011
Earlier work this paper cites.
Computation and palaeography: potentials and limits
Tal Hassner, Malte Rehbein, Peter A Stokes, and Lior Wolf · 2012
Earlier work this paper cites.
Scene text recognition using higher order language priors
Anand Mishra, Karteek Alahari, and C.V. Jawahar · 2012
Earlier work this paper cites.
Detecting texts of arbitrary orientations in natural images
Cong Yao, Xiang Bai, Wenyu Liu, Yi Ma, and Zhuowen Tu · 2012
Earlier work this paper cites.
Icdar 2013 robust reading competition
Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluis Gomez i Bigorda, Sergi Robles Mestre, Joan Mas, David Fernandez Mota, Jon Almazan Almazan, and Lluis Pere De Las Heras · 2013
Earlier work this paper cites.
Recognizing text with perspec- tive distortion in natural scenes
Trung Quy Phan, Palaiahnakote Shivakumara, Shangxuan Tian, and Chew Lim Tan · 2013
Earlier work this paper cites.
Age and gender estimation of unfiltered faces
Eran Eidinger, Roee Enbar, and Tal Hassner · 2014
Earlier work this paper cites.
Digital palaeography: New machines and old texts (dagstuhl seminar 14302)
Tal Hassner, Robert Sablatnig, Dominique Stutzmann, and Ségolène Tarte · 2014
Earlier work this paper cites.
Synthetic data and artificial neural networks for natural scene text recognition
Max Jaderberg, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Recognizing text with perspec- tive distortion in natural scenes
Anhar Risnumawan, Palaiahankote Shivakumara, Chee Seng Chan, and Chew Lim Tan · 2014
Earlier work this paper cites.
Icdar 2015 competition on robust reading
Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman Ghosh, Andrew Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ramaseshan Chandrasekhar, Shijian Lu, et al · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross B. Girshick, and J. Sun · 2015
Earlier work this paper cites.
Cider: Consensus-based image description evaluation
Ramakrishna Vedantam, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Synthetic data for text localisation in natural images
Ankush Gupta, Andrea Vedaldi, and Andrew Zisserman · 2016
Cited alongside, same era.
Fasttext. zip: Compressing text classification models
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov · 2016
Cited alongside, same era.
Star-net: A spatial attention residue network for scene text recognition
Wei Liu, Chaofeng Chen, Kwan-Yee K. Wong, Zhizhong Su, and Junyu Han · 2016
Cited alongside, same era.
Coco-text: Dataset and benchmark for text detection and recognition in natural images
Andreas Veit, Tomas Matera, Lukas Neumann, Jiri Matas, and Serge Belongie · 2016
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Later among the works it cites.
Towards unconstrained end-to-end text spotting
Siyang Qin, Alessandro Bissaco, Michalis Raptis, Yasuhisa Fujii, and Ying Xiao · 2019
Later among the works it cites.
Towards vqa models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach · 2019
Later among the works it cites.
Icdar 2019 competition on large-scale street view text with partial labeling – rrc-lsvt
Yipeng Sun, Zihan Ni, Chee-Kheng Chng, Yuliang Liu, Canjie Luo, Chun Chet Ng, Junyu Han, Errui Ding, Jingtuo Liu, Dimosthenis Karatzas, Chee Seng Chan, and Lianwen Jin · 2019
Later among the works it cites.
Leaf-qa: Locate, encode & attend for figure question answering
Ritwick Chaudhry, Sumit Shekhar, Utkarsh Gupta, Pranav Maneriker, Prann Bansal, and Ajay Joshi · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Detecting curve text in the wild: New dataset and new solution
Yuliang Liu, Lianwen Jin, Shuaitao Zhang, and Sheng Zhang · 2017
Cited alongside, same era.
Icdar2017 robust reading challenge on multi-lingual scene text detection and script identification - rrc-mlt
N. Nayef, F. Yin, I. Bizid, H. Choi, Y. Feng, D. Karatzas, Z. Luo, U. Pal, C. Rigaud, J. Chazalon, W. Khlif, M. M. Luqman, J.-C. Burie, C.-L. Liu, and J.-M. Ogier · 2017
Cited alongside, same era.
An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition
Baoguang Shi, Xiang Bai, and Cong Yao · 2017
Cited alongside, same era.
Icdar2017 competition on reading chinese text in the wild (rctw-17)
Baoguang Shi, Cong Yao, Minghui Liao, Mingkun Yang, Pei Xu, Linyan Cui, Serge Belongie, Shijian Lu, and Xiang Bai · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Cited alongside, same era.
Rosetta: Large scale system for text detection and recognition in images
Fedor Borisyuk, Albert Gordo, and Viswanath Sivakumar · 2018
Cited alongside, same era.
Rosetta: Large scale system for text detection and recognition in images
Fedor Borisyuk, Albert Gordo, and Viswanath Sivakumar · 2018
Cited alongside, same era.
Later among the works it cites.
Structured multimodal attentions for textvqa
Chenyu Gao, Qi Zhu, Peng Wang, Hui Li, Y. Liu, A. V. D. Hengel, and Qi Wu · 2020
Later among the works it cites.
Captioning images taken by people who are blind
Danna Gurari, Yinan Zhao, Meng Zhang, and Nilavra Bhattacharya · 2020
Later among the works it cites.
Finding the evidence: Localization-aware answer prediction for text visual question answering
Wei Han, Hantao Huang, and T. Han · 2020
Later among the works it cites.
Iterative answer prediction with pointer-augmented multimodal transformers for textvqa
Ronghang Hu, Amanpreet Singh, Trevor Darrell, and Marcus Rohrbach · 2020
Later among the works it cites.
Ruart: A novel text-centered solution for text-based visual question answering
Zan-Xia Jin, Heran Wu, C. Yang, Fang Zhou, Jingyan Qin, Lei Xiao, and XuCheng Yin · 2020
Later among the works it cites.
Spatially aware multimodal transformers for textvqa
Yash Kant, Dhruv Batra, Peter Anderson, A. Schwing, D. Parikh, Jiasen Lu, and Harsh Agrawal · 2020
Later among the works it cites.
The hateful memes challenge: Detecting hate speech in multimodal memes
Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine · 2020
Later among the works it cites.
Mask textspotter v3: Segmentation proposalnetwork for robust scene text spotting
Minghui Liao, Guan Pang, Jing Huang, Tal Hassner, and Xiang Bai · 2020
Later among the works it cites.
Abcnet: Real-time scene text spotting with adaptive bezier-curve network
Yuliang Liu*, Hao Chen*, Chunhua Shen, Tong He, Lianwen Jin, and Liangwei Wang · 2020
Later among the works it cites.
DocVQA: A dataset for vqa on document images, 2020
Minesh Mathew, Dimosthenis Karatzas, R. Manmatha, and C. V. Jawahar · 2020
Later among the works it cites.
Plotqa: Reasoning over scientific plots
Nitesh Methani, Pritha Ganguly, Mitesh M Khapra, and Pratyush Kumar · 2020
Later among the works it cites.
Textcaps: a dataset for image captioning with reading comprehension
Oleksii Sidorov, Ronghang Hu, Marcus Rohrbach, and Amanpreet Singh · 2020
Later among the works it cites.
Mmf: A multimodal framework for vision and language research
Amanpreet Singh, Vedanuj Goswami, Vivek Natarajan, Yu Jiang, Xinlei Chen, Meet Shah, Marcus Rohrbach, Dhruv Batra, and Devi Parikh · 2020
Later among the works it cites.
All you need is boundary: Toward arbitrary-shaped text spotting
Hao Wang*, Pu Lu*, Hui Zhang*, Mingkun Yang, Xiang Bai, Yongchao Xu, Mengchao He, Yongpan Wang, and Wenyu Liu · 2020
Later among the works it cites.
On the general value of evidence, and bilingual scene-text visual question answering
Xinyu Wang, Yuliang Liu, Chunhua Shen, Chun Chet Ng, Canjie Luo, Lianwen Jin, Chee Seng Chan, Anton van den Hengel, and Liangwei Wang · 2020
Later among the works it cites.
A multiplexed network for end-to-end, multilingual ocr
Jing Huang, Guan Pang, Rama Kovvuri, Mandy Toh, Kevin J Liang, Praveen Krishnan, Xi Yin, and Tal Hassner · 2021
Closest in time.