Fetching the paper…
Reading the bibliography…
In order to answer semantically-complicated questions about an image, a Visual Question Answering (VQA) model needs to fully understand the visual scene in the image, especially the interactive dynamics between different objects.
Scene perception: Detecting and judging objects undergoing relational violations
Irving Biederman, Robert J Mezzanotte, and Jan C Rabinowitz · 1982
Earlier work this paper cites.
Object categorization using co-occurrence, location and appearance
Carolina Galleguillos, Andrew Rabinovich, and Serge Belongie · 2008
Earlier work this paper cites.
Multi-class segmentation with relative location prior
Stephen Gould, Jim Rodgers, David Cohen, Gal Elidan, and Daphne Koller · 2008
Earlier work this paper cites.
An empirical study of context in object detection
Santosh K Divvala, Derek Hoiem, James H Hays, Alexei A Efros, and Martial Hebert · 2009
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
Ali Farhadi, Mohsen Hejrati, Mohammad Amin Sadeghi, Peter Young, Cyrus Rashtchian, Julia Hockenmaier, and David Forsyth · 2010
Earlier work this paper cites.
Object detection with discriminatively trained part-based models
Pedro F Felzenszwalb, Ross B Girshick, David McAllester, and Deva Ramanan · 2010
Earlier work this paper cites.
Recognition using visual phrases
Mohammad Amin Sadeghi and Ali Farhadi · 2011
Earlier work this paper cites.
A tree-based context model for object recognition
Myung Jin Choi, Antonio Torralba, and Alan S Willsky · 2012
Earlier work this paper cites.
Learning everything about anything: Webly-supervised visual concept learning
Santosh K Divvala, Ali Farhadi, and Carlos Guestrin · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick · 2014
Earlier work this paper cites.
The role of context for object detection and semantic segmentation in the wild
Roozbeh Mottaghi, Xianjie Chen, Xiaobai Liu, Nam-Gyu Cho, Seong-Whan Lee, Sanja Fidler, Raquel Urtasun, and Alan Yuille · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
From captions to visual concepts and back
Hao Fang, Saurabh Gupta, Forrest Iandola, Rupesh K Srivastava, Li Deng, Piotr Dollár, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John C Platt, et al · 2015
Earlier work this paper cites.
Image retrieval using scene graphs
Justin Johnson, Ranjay Krishna, Michael Stark, Li-Jia Li, David Shamma, Michael Bernstein, and Li Fei-Fei · 2015
Earlier work this paper cites.
Learning semantic relationships for better action retrieval in images
Vignesh Ramanathan, Congcong Li, Jia Deng, Wei Han, Zhen Li, Kunlong Gu, Yang Song, Samy Bengio, Charles Rosenberg, and Li Fei-Fei · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Generating semantically precise scene graphs from textual descriptions for improved image retrieval
Sebastian Schuster, Ranjay Krishna, Angel Chang, Li Fei-Fei, and Christopher D Manning · 2015
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, Michael Bernstein, and Li Fei-Fei · 2016
Earlier work this paper cites.
Visual relationship detection with language priors
Cewu Lu, Ranjay Krishna, Michael Bernstein, and Li Fei-Fei · 2016
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2016
Cited alongside, same era.
Ask me anything: Free-form visual question answering based on knowledge from external sources
Qi Wu, Peng Wang, Chunhua Shen, Anthony Dick, and Anton van den Hengel · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola · 2016
Cited alongside, same era.
Mutan: Multimodal tucker fusion for visual question answering
Hedi Ben-Younes, Rémi Cadene, Matthieu Cord, and Nicolas Thome · 2017
Cited alongside, same era.
Detecting visual relationships with deep relational networks
Bo Dai, Yuqi Zhang, and Dahua Lin · 2017
Don’t just assume; look and answer: Overcoming priors for visual question answering
Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi · 2018
Later among the works it cites.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Later among the works it cites.
Stacked latent attention for multimodal reasoning
Haoqi Fan and Jiatong Zhou · 2018
Later among the works it cites.
Relation networks for object detection
Han Hu, Jiayuan Gu, Zheng Zhang, Jifeng Dai, and Yichen Wei · 2018
Later among the works it cites.
Pythia v0. 1: the winning entry to the vqa challenge 2018
Yu Jiang, Vivek Natarajan, Xinlei Chen, Marcus Rohrbach, Dhruv Batra, and Devi Parikh · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Accurate, large minibatch sgd: training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Cited alongside, same era.
Hadamard product for low-rank bilinear pooling
Jin-Hwa Kim, Kyoung-Woon On, Woosang Lim, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang · 2017
Cited alongside, same era.
Semi-supervised classification with graph convolutional networks
Thomas N Kipf and Max Welling · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2017
Cited alongside, same era.
Dual attention networks for multimodal reasoning and matching
Hyeonseob Nam, Jung-Woo Ha, and Jeonghee Kim · 2017
Cited alongside, same era.
Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang · 2018
Later among the works it cites.
Incorporating external knowledge to answer open-domain visual questions with dynamic memory networks
Guohao Li, Hang Su, and Wenwu Zhu · 2018
Later among the works it cites.
Tell-and-answer: Towards explainable visual question answering using attributes and captions
Qing Li, Jianlong Fu, Dongfei Yu, Tao Mei, and Jiebo Luo · 2018
Later among the works it cites.
R-vqa: Learning visual relation facts with semantic attention for visual question answering
Pan Lu, Lei Ji, Wei Zhang, Nan Duan, Ming Zhou, and Jianyong Wang · 2018
Later among the works it cites.
Visual question answering with memory-augmented networks
Chao Ma, Chunhua Shen, Anthony Dick, Qi Wu, Peng Wang, Anton van den Hengel, and Ian Reid · 2018
Later among the works it cites.
Learning visual question answering by bootstrapping hard attention
Mateusz Malinowski, Carl Doersch, Adam Santoro, and Peter Battaglia · 2018
Later among the works it cites.
Learning conditioned graph structures for interpretable visual question answering
Will Norcliffe-Brown, Stathis Vafeias, and Sarah Parisot · 2018
Later among the works it cites.
Differential attention for visual question answering
Badri Patro and Vinay P Namboodiri · 2018
Later among the works it cites.
Learning visual knowledge memory networks for visual question answering
Zhou Su, Chen Zhu, Yinpeng Dong, Dongqi Cai, Yurong Chen, and Jianguo Li · 2018
Later among the works it cites.
Tips and tricks for visual question answering: Learnings from the 2017 challenge
Damien Teney, Peter Anderson, Xiaodong He, and Anton van den Hengel · 2018
Later among the works it cites.
Multi-modal learning with prior visual relation reasoning
Zhuoqian Yang, Jing Yu, Chenghao Yang, Zengchang Qin, and Yue Hu · 2018
Later among the works it cites.
Exploring visual relationship for image captioning
Ting Yao, Yingwei Pan, Yehao Li, and Tao Mei · 2018
Later among the works it cites.
Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering
Zhou Yu, Jun Yu, Chenchao Xiang, Jianping Fan, and Dacheng Tao · 2018
Later among the works it cites.
Learning to count objects in natural images for visual question answering
Yan Zhang, Jonathon S. Hare, and Adam Prügel-Bennett · 2018
Later among the works it cites.
Murel: Multimodal relational reasoning for visual question answering
Remi Cadene, Hedi Ben-younes, Matthieu Cord, and Nicolas Thome · 2019
Closest in time.
Neighbourhood watch: Referring expression comprehension via language-guided graph attention networks
Peng Wang, Qi Wu, Jiewei Cao, Chunhua Shen, Lianli Gao, and Anton van den Hengel · 2019
Closest in time.