Fetching the paper…
Reading the bibliography…
Modern Visual Question Answering (VQA) models have been shown to rely heavily on superficial correlations between question and answer words learned during training such as overwhelmingly reporting the type of room as kitchen or the sport being played as tennis, irrespective of the image.
Unbiased look at dataset bias
Antonio Torralba and Alexei A Efros · 2011
Earlier work this paper cites.
Reporting bias and knowledge acquisition
Jonathan Gordon and Benjamin Van Durme · 2013
Earlier work this paper cites.
Generative adversarial networks, 2014
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio · 2014
Earlier work this paper cites.
Conditional generative adversarial nets
Mehdi Mirza and Simon Osindero · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Neural module networks, 2015
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein · 2015
Earlier work this paper cites.
VQA: Visual Question Answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Unsupervised representation learning with deep convolutional generative adversarial networks, 2015
Alec Radford, Luke Metz, and Soumith Chintala · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Stacked attention networks for image question answering, 2015
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola · 2015
Earlier work this paper cites.
Analyzing the behavior of visual question answering models
Aishwarya Agrawal, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering, 2016
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2016
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering, 2016
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh · 2016
Cited alongside, same era.
Yin and yang: Balancing and answering binary visual questions
Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2016
Cited alongside, same era.
Visual7w: Grounded question answering in images
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei · 2016
Cited alongside, same era.
Don’t just assume; look and answer: Overcoming priors for visual question answering, 2017
Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi · 2017
Cited alongside, same era.
Inferring and executing programs for visual reasoning, 2017
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Judy Hoffman, Li Fei-Fei, C. Lawrence Zitnick, and Ross Girshick · 2017
Later among the works it cites.
An analysis of visual question answering algorithms
Kushal Kafle and Christopher Kanan · 2017
Later among the works it cites.
Fader networks: Manipulating images by sliding attributes, 2017
Guillaume Lample, Neil Zeghidour, Nicolas Usunier, Antoine Bordes, Ludovic Denoyer, and Marc’Aurelio Ranzato · 2017
Later among the works it cites.
Learning to pivot with adversarial networks
Gilles Louppe, Michael Kagan, and Kyle Cranmer · 2017
Later among the works it cites.
Learning visual reasoning without strong priors, 2017
Ethan Perez, Harm de Vries, Florian Strub, Vincent Dumoulin, and Aaron Courville · 2017
Later among the works it cites.
Adversarial discriminative domain adaptation
Eric Tzeng, Judy Hoffman, Trevor Darrell, and Kate Saenko · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aishwarya Agrawal, Aniruddha Kembhavi, Dhruv Batra, and Devi Parikh · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering, 2017
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2017
Cited alongside, same era.
Towards diverse and natural image descriptions via a conditional gan
Bo Dai, Sanja Fidler, Raquel Urtasun, and Dahua Lin · 2017
Cited alongside, same era.
Visual Dialog
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José M.F. Moura, Devi Parikh, and Dhruv Batra · 2017
Cited alongside, same era.
Learning to reason: End-to-end module networks for visual question answering, 2017
Ronghang Hu, Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Kate Saenko · 2017
Cited alongside, same era.
Clevr: A diagnostic dataset for compositional language and elementary visual reasoning
Justin Johnson, Bharath Hariharan, Laurens van der Maaten, Li Fei-Fei, C Lawrence Zitnick, and Ross Girshick · 2017
Cited alongside, same era.
Later among the works it cites.
Stackgan: Text to photo-realistic image synthesis with stacked generative adversarial networks
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaolei Huang, Xiaogang Wang, and Dimitris Metaxas · 2017
Later among the works it cites.
Men also like shopping: Reducing gender bias amplification using corpus-level constraints, 2017
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang · 2017
Later among the works it cites.
Women also snowboard: Overcoming bias in captioning models
Kaylee Burns, Lisa Anne Hendricks, Trevor Darrell, and Anna Rohrbach · 2018
Closest in time.
Vizwiz grand challenge: Answering visual questions from blind people
Danna Gurari, Qing Li, Abigale J Stangl, Anhong Guo, Chi Lin, Kristen Grauman, Jiebo Luo, and Jeffrey P Bigham · 2018
Closest in time.