Fetching the paper…
Reading the bibliography…
Visual Question Answering (VQA) has received a lot of attention over the past couple of years.
Relevance feedback in image retrieval: A comprehensive review
Xiang Sean Zhou and Thomas S. Huang · 2003
Earlier work this paper cites.
Learning to detect unseen object classes by between-class attribute transfer
Christoph H Lampert, Hannes Nickisch, and Stefan Harmeling · 2009
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
Mateusz Malinowski and Mario Fritz · 2014
Earlier work this paper cites.
Decorrelating semantic visual attributes by resisting the urge to share
Dinesh Jayaraman, Fei Sha, and Kristen Grauman · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Visual turing test for computer vision systems
Donald Geman, Stuart Geman, Neil Hallonquist, and Laurent Younes · 2015
Earlier work this paper cites.
Are you talking to a machine? dataset and methods for multilingual image question
Haoyuan Gao, Junhua Mao, Jie Zhou, Zhiheng Huang, Lei Wang, and Wei Xu · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
Mengye Ren, Ryan Kiros, and Richard Zemel · 2015
Earlier work this paper cites.
ABC-CNN: an attention based convolutional neural network for visual question answering
Kan Chen, Jiang Wang, Liang-Chieh Chen, Haoyuan Gao, Wei Xu, and Ram Nevatia · 2015
Earlier work this paper cites.
Compositional memory for visual question answering
Aiwen Jiang, Fang Wang, Fatih Porikli, and Yi Li · 2015
Earlier work this paper cites.
Explicit knowledge-based reasoning for visual question answering
Peng Wang, Qi Wu, Chunhua Shen, Anton van den Hengel, and Anthony R. Dick · 2015
Earlier work this paper cites.
Deeper lstm and normalized cnn visual question answering model
Jiasen Lu, Xiao Lin, Dhruv Batra, and Devi Parikh · 2015
Earlier work this paper cites.
Analyzing the behavior of visual question answering models
Aishwarya Agrawal, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Yin and Yang: Balancing and answering binary visual questions
Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2016
Cited alongside, same era.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2016
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2016
Cited alongside, same era.
Guesswhat?! visual object discovery through multi-modal dialogue
Harm de Vries, Florian Strub, Sarath Chandar, Olivier Pietquin, Hugo Larochelle, and Aaron Courville · 2016
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Multimodal residual learning for visual QA
Jin-Hwa Kim, Sang-Woo Lee, Dong-Hyun Kwak, Min-Oh Heo, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang · 2016
Later among the works it cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Later among the works it cites.
Training recurrent answering units with joint loss minimization for vqa
Hyeonwoo Noh and Bohyung Han · 2016
Later among the works it cites.
A focused dynamic attention model for visual question answering
Ilija Ilievski, Shuicheng Yan, and Jiashi Feng · 2016
Later among the works it cites.
Ask me anything: Free-form visual question answering based on knowledge from external sources
Qi Wu, Peng Wang, Chunhua Shen, Anton van den Hengel, and Anthony R. Dick · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2016
Cited alongside, same era.
Visual7w: Grounded question answering in images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alexander J. Smola · 2016
Cited alongside, same era.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
Huijuan Xu and Kate Saenko · 2016
Cited alongside, same era.
Deep compositional question answering with neural module networks
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Cited alongside, same era.
Answer-type prediction for visual question answering
Kushal Kafle and Christopher Kanan · 2016
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2016
Cited alongside, same era.
Learning to compose neural networks for question answering
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2016
Cited alongside, same era.
Dynamic memory networks for visual and textual question answering
Caiming Xiong, Stephen Merity, and Richard Socher · 2016
Later among the works it cites.
Dualnet: Domain-invariant network for visual question answering
Kuniaki Saito, Andrew Shin, Yoshitaka Ushiku, and Tatsuya Harada · 2016
Later among the works it cites.
Learning to generalize to new compositions in image understanding
Yuval Atzmon, Jonathan Berant, Vahid Kezami, Amir Globerson, and Gal Chechik · 2016
Later among the works it cites.
Zero-shot visual question answering
Damien Teney and Anton van den Hengel · 2016
Later among the works it cites.
Visual Dialog
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José M.F. Moura, Devi Parikh, and Dhruv Batra · 2017
Closest in time.
Learning cooperative visual dialog agents with deep reinforcement learning
Abhishek Das, Satwik Kottur, José M.F. Moura, Stefan Lee, and Dhruv Batra · 2017
Closest in time.
Image-grounded conversations: Multimodal context for natural question and response generation
Nasrin Mostafazadeh, Chris Brockett, Bill Dolan, Michel Galley, Jianfeng Gao, Georgios P. Spithourakis, and Lucy Vanderwende · 2017
Closest in time.
An empirical evaluation of visual question answering for novel objects
Santhosh K Ramakrishnan, Ambar Pal, Gaurav Sharma, and Anurag Mittal · 2017
Closest in time.