Fetching the paper…
Reading the bibliography…
Benefiting from the advancement of computer vision, natural language processing and information retrieval techniques, visual question answering (VQA), which aims to answer questions about an image or a video, has received lots of attentions over the past few years.
Analyzing the behavior of visual question answering models. In EMNLP . ACL, 1955–1960
Aishwarya Agrawal, Dhruv Batra, and Devi Parikh. 2016 · 1960
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
CLEVR: A diagnostic dataset for compositional language and elementary visual reasoning. In CVPR . IEEE, 1988–1997
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick. 2017 · 1997
Earlier work this paper cites.
High performance question/answering. In SIGIR . ACM, 366–374
Marius A Pasca and Sandra M Harabagiu. 2001 · 2001
Earlier work this paper cites.
Multimedia answering: enriching text QA with media information. In SIGIR . ACM, 695–704
Liqiang Nie, Meng Wang, Zhengjun Zha, Guangda Li, and Tat-Seng Chua. 2011 · 2011
Earlier work this paper cites.
Microsoft coco: Common objects in context. In ECCV . Springer, 740–755
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input. In NIPS . 1682–1690
Mateusz Malinowski and Mario Fritz. 2014 · 2014
Earlier work this paper cites.
Two-stream convolutional networks for action recognition in videos. In NIPS . 568–576
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering. In ICCV . IEEE, 2425–2433
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions. In CVPR . IEEE, 3128–3137
Andrej Karpathy and Li Fei-Fei. 2015 · 2015
Earlier work this paper cites.
Ask your neurons: A neural-based approach to answering questions about images. In ICCV . IEEE, 1–9
Mateusz Malinowski, Marcus Rohrbach, and Mario Fritz. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks. In NIPS . 91–99
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition. In ICLR
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Earlier work this paper cites.
Neural module networks. In CVPR . IEEE, 39–48
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In CVPR . IEEE, 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Composite correlation quantization for efficient multimodal retrieval. In SIGIR . ACM, 579–588
Mingsheng Long, Yue Cao, Jianmin Wang, and Philip S Yu. 2016 · 2016
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering. In NIPS . 289–297
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2016 · 2016
Cited alongside, same era.
Novelty based ranking of human answers for community questions. In SIGIR . ACM, 215–224
Adi Omari, David Carmel, Oleg Rokhlenko, and Idan Szpektor. 2016 · 2016
Cited alongside, same era.
Deep dynamic neural networks for multimodal gesture segmentation and recognition
Di Wu, Lionel Pigou, Pieter-Jan Kindermans, Nam Do-Hoang Le, Ling Shao, Joni Dambre, and Jean-Marc Odobez. 2016a · 2016
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata andJoshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. 2017 · 2017
Later among the works it cites.
Deep multimodal learning: A survey on recent advances and trends
Dhanesh Ramachandram and Graham W Taylor. 2017 · 2017
Later among the works it cites.
Visual question answering: A survey of methods and datasets
Qi Wu, Damien Teney, Peng Wang, Chunhua Shen, Anthony Dick, and Anton van den Hengel. 2017 · 2017
Later among the works it cites.
Multi-modal factorized bilinear pooling with co-attention learning for visual question answering. In ICCV . IEEE, 1839–1848
Zhou Yu, Jun Yu, Jianping Fan, and Dacheng Tao. 2017 · 2017
Later among the works it cites.
Learning max-margin geoSocial multimedia network representations for point-of-interest suggestion. In SIGIR . ACM, 833–836
Zhou Zhao, Qifan Yang, Hanqing Lu, Min Yang, Jun Xiao, Fei Wu, and Yueting Zhuang. 2017 · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering. In ECCV . Springer, 451–466
Huijuan Xu and Kate Saenko. 2016 · 2016
Cited alongside, same era.
Stacked attention networks for image question answering. In CVPR . IEEE, 21–29
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola. 2016 · 2016
Cited alongside, same era.
Yin and yang: Balancing and answering binary visual questions. In CVPR . IEEE, 5014–5022
Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2016 · 2016
Cited alongside, same era.
Visual7w: Grounded question answering in images. In CVPR . IEEE, 4995–5004
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei. 2016 · 2016
Cited alongside, same era.
Attentive collaborative filtering: Multimedia recommendation with item-and component-level attention. In SIGIR . ACM, 335–344
Jingyuan Chen, Hanwang Zhang, Xiangnan He, Liqiang Nie, Wei Liu, and Tat-Seng Chua. 2017 · 2017
Cited alongside, same era.
A hierarchical multimodal attention-based neural network for image captioning. In SIGIR . ACM, 889–892
Yong Cheng, Fei Huang, Lian Zhou, Cheng Jin, Yuejie Zhang, and Tao Zhang. 2017 · 2017
Cited alongside, same era.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In CVPR . IEEE, 6325–6334
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Don’t just assume; look and answer: Overcoming priors for visual question answering. In CVPR . IEEE, 4971–4980
Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi. 2018 · 2018
Later among the works it cites.
Bottom-up and top-down attention for image captioning and visual question answering. In CVPR . IEEE, 6077–6086
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Later among the works it cites.
Deep attention neural tensor network for visual question answering. In ECCV . Springer, 21–37
Yalong Bai, Jianlong Fu, Tiejun Zhao, and Tao Mei. 2018 · 2018
Later among the works it cites.
User profiling through deep multimodal fusion. In WSDM . ACM, 171–179
Golnoosh Farnadi, Jie Tang, Martine De Cock, and Marie-Francine Moens. 2018 · 2018
Later among the works it cites.
Multi-modal preference modeling for product search. In MM . ACM, 1865–1873
Yangyang Guo, Zhiyong Cheng, Liqiang Nie, Xin-Shun Xu, and Mohan Kankanhalli. 2018 · 2018
Later among the works it cites.
Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering. In CVPR . IEEE, 6087–6096
Duy-Kien Nguyen and Takayuki Okatani. 2018 · 2018
Later among the works it cites.
Overcoming language priors in visual question answering with adversarial regularization. In NIPS . 1546–1556
Sainandan Ramakrishnan, Aishwarya Agrawal, and Stefan Lee. 2018 · 2018
Later among the works it cites.
Learning to count objects in natural images for visual question answering. In ICLR
Yan Zhang, Jonathon Hare, and Adam Prügel-Bennett. 2018 · 2018
Later among the works it cites.
MMALFM: Explainable recommendation by leveraging reviews and images
Zhiyong Cheng, Xiaojun Chang, Lei Zhu, Rose C Kanjirathinkal, and Mohan Kankanhalli. 2019 · 2019
Closest in time.