Fetching the paper…
Reading the bibliography…
Current visual question answering datasets do not consider the rich semantic information conveyed by text within an image.
Conditioned reflex: An investigation of the physiological activity of the cerebral cortex
Ivan P Pavlov · 1960
Earlier work this paper cites.
Operant behavior
Burrhus F Skinner · 1963
Earlier work this paper cites.
Binary codes capable of correcting deletions, insertions, and reversals
Vladimir I Levenshtein · 1966
Earlier work this paper cites.
Document image retrieval for QA systems based on the density distributions of successive terms
Koichi Kise, Shota Fukushima, and Keinosuke Matsumoto · 2005
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Word spotting in the wild
Kai Wang and Serge Belongie · 2010
Earlier work this paper cites.
Undoing the damage of dataset bias
Aditya Khosla, Tinghui Zhou, Tomasz Malisiewicz, Alexei A Efros, and Antonio Torralba · 2012
Earlier work this paper cites.
Icdar 2013 robust reading competition
Dimosthenis Karatzas, Faisal Shafait, Seiichi Uchida, Masakazu Iwamura, Lluis Gomez i Bigorda, Sergi Robles Mestre, Joan Mas, David Fernandez Mota, Jon Almazan Almazan, and Lluis Pere De Las Heras · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Earlier work this paper cites.
Image retrieval using textual cues
A. Mishra, K. Alahari, and C. V. Jawahar · 2013
Earlier work this paper cites.
Word spotting and recognition with embedded attributes
Jon Almazán, Albert Gordo, Alicia Fornés, and Ernest Valveny · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
Mateusz Malinowski and Mario Fritz · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Are you talking to a machine? dataset and methods for multilingual image question
Haoyuan Gao, Junhua Mao, Jie Zhou, Zhiheng Huang, Lei Wang, and Wei Xu · 2015
Earlier work this paper cites.
Icdar 2015 competition on robust reading
Dimosthenis Karatzas, Lluis Gomez-Bigorda, Anguelos Nicolaou, Suman Ghosh, Andrew Bagdanov, Masakazu Iwamura, Jiri Matas, Lukas Neumann, Vijay Ramaseshan Chandrasekhar, Shijian Lu, et al · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
Mengye Ren, Ryan Kiros, and Richard Zemel · 2015
Earlier work this paper cites.
Pointer networks
Oriol Vinyals, Meire Fortunato, and Navdeep Jaitly · 2015
Cited alongside, same era.
Multi-orientation scene text detection with adaptive clustering
Xu-Cheng Yin, Wei-Yi Pei, Jun Zhang, and Hong-Wei Hao · 2015
Cited alongside, same era.
Visual madlibs: Fill in the blank description generation and question answering
Licheng Yu, Eunbyung Park, Alexander C Berg, and Tamara L Berg · 2015
Cited alongside, same era.
Pointing the unknown words
Caglar Gulcehre, Sungjin Ahn, Ramesh Nallapati, Bowen Zhou, and Yoshua Bengio · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Reading text in the wild with convolutional neural networks
Max Jaderberg, Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman · 2016
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
J. Johnson, B. Hariharan, L. van der Maaten, L. Fei-Fei, C. Lawrence Zitnick, and R. Girshick · 2017
Later among the works it cites.
Figureqa: An annotated figure dataset for visual reasoning
Samira Ebrahimi Kahou, Vincent Michalski, Adam Atkinson, Ákos Kádár, Adam Trischler, and Yoshua Bengio · 2017
Later among the works it cites.
Show, ask, attend, and answer: A strong baseline for visual question answering
Vahid Kazemi and Ali Elqursh · 2017
Later among the works it cites.
Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
Aniruddha Kembhavi, Minjoon Seo, Dustin Schwenk, Jonghyun Choi, Ali Farhadi, and Hannaneh Hajishirzi · 2017
Later among the works it cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dynamic lexicon generation for natural scene images
Yash Patel, Lluis Gomez, Marçal Rusinol, and Dimosthenis Karatzas · 2016
Cited alongside, same era.
Robust scene text recognition with automatic rectification
Baoguang Shi, Xinggang Wang, Pengyuan Lyu, Cong Yao, and Xiang Bai · 2016
Cited alongside, same era.
Movieqa: Understanding stories in movies through question-answering
Makarand Tapaswi, Yukun Zhu, Rainer Stiefelhagen, Antonio Torralba, Raquel Urtasun, and Sanja Fidler · 2016
Cited alongside, same era.
Coco-text: Dataset and benchmark for text detection and recognition in natural images
Andreas Veit, Tomas Matera, Lukas Neumann, Jiri Matas, and Serge Belongie · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola · 2016
Cited alongside, same era.
Yin and yang: Balancing and answering binary visual questions
Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2016
Cited alongside, same era.
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Later among the works it cites.
Textboxes: A fast text detector with a single deep neural network
Minghui Liao, Baoguang Shi, Xiang Bai, Xinggang Wang, and Wenyu Liu · 2017
Later among the works it cites.
East: an efficient and accurate scene text detector
Xinyu Zhou, Cong Yao, He Wen, Yuzhi Wang, Shuchang Zhou, Weiran He, and Jiajun Liang · 2017
Later among the works it cites.
Don’t just assume; look and answer: Overcoming priors for visual question answering
Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi · 2018
Later among the works it cites.
Rosetta: Large scale system for text detection and recognition in images
Fedor Borisyuk, Albert Gordo, and Viswanath Sivakumar · 2018
Later among the works it cites.
Single shot scene text retrieval
Lluís Gómez, Andrés Mafla, Marçal Rusinol, and Dimosthenis Karatzas · 2018
Later among the works it cites.
Vizwiz grand challenge: Answering visual questions from blind people
Danna Gurari, Qing Li, Abigale J Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P Bigham · 2018
Later among the works it cites.
An end-to-end textspotter with explicit alignment and attention
Tong He, Zhi Tian, Weilin Huang, Chunhua Shen, Yu Qiao, and Changming Sun · 2018
Later among the works it cites.
Dvqa: Understanding data visualizations via question answering
Kushal Kafle, Brian Price, Scott Cohen, and Christopher Kanan · 2018
Later among the works it cites.
Textboxes++: A single-shot oriented scene text detector
Minghui Liao, Baoguang Shi, and Xiang Bai · 2018
Later among the works it cites.
Fots: Fast oriented text spotting with a unified network
Xuebo Liu, Ding Liang, Shi Yan, Dagui Chen, Yu Qiao, and Junjie Yan · 2018
Later among the works it cites.
Mask textspotter: An end-to-end trainable neural network for spotting text with arbitrary shapes
Pengyuan Lyu, Minghui Liao, Cong Yao, Wenhao Wu, and Xiang Bai · 2018
Later among the works it cites.
Pythia-a platform for vision & language research
Amanpreet Singh, Vivek Natarajan, Yu Jiang, Xinlei Chen, Meet Shah, Marcus Rohrbach, Dhruv Batra, and Devi Parikh · 2018
Later among the works it cites.
Tallyqa: Answering complex counting questions
Manoj Acharya, Kushal Kafle, and Christopher Kanan · 2019
Closest in time.
Icdar 2019 competition on scene text visual question answering
Ali Furkan Biten, Rubèn Tito, Andres Mafla, Lluis Gomez, Marçal Rusiñol, Minesh Mathew, CV Jawahar, Ernest Valveny, and Dimosthenis Karatzas · 2019
Closest in time.
Towards vqa models that can read
Amanpreet Singh, Vivek Natarajan, Meet Shah, Yu Jiang, Xinlei Chen, Dhruv Batra, Devi Parikh, and Marcus Rohrbach · 2019
Closest in time.