Fetching the paper…
Reading the bibliography…
While Visual Question Answering (VQA) has progressed rapidly, previous works raise concerns about robustness of current VQA models.
Context-based vision system for place and object recognition
Antonio Torralba, Kevin P Murphy, William T Freeman, and Mark A Rubin · 2003
Earlier work this paper cites.
The hadamard product
Elizabeth Million · 2007
Earlier work this paper cites.
The role of context in object recognition
Aude Oliva and Antonio Torralba · 2007
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
Mateusz Malinowski and Mario Fritz · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh · 2015
Earlier work this paper cites.
Abc-cnn: An attention based convolutional neural network for visual question answering
Kan Chen, Jiang Wang, Liang-Chieh Chen, Haoyuan Gao, Wei Xu, and Ram Nevatia · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun · 2015
Earlier work this paper cites.
Simple baseline for visual question answering
Bolei Zhou, Yuandong Tian, Sainbayar Sukhbaatar, Arthur Szlam, and Rob Fergus · 2015
Earlier work this paper cites.
Analyzing the behavior of visual question answering models
Aishwarya Agrawal, Dhruv Batra, and Devi Parikh · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach · 2016
Earlier work this paper cites.
Hadamard product for low-rank bilinear pooling
Jin-Hwa Kim, Kyoung-Woon On, Woosang Lim, Jeonghee Kim, Jung-Woo Ha, and Byoung-Tak Zhang · 2016
Earlier work this paper cites.
Stacked attention networks for image question answering
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola · 2016
Earlier work this paper cites.
Mutan: Multimodal tucker fusion for visual question answering
Hedi Ben-Younes, Rémi Cadene, Matthieu Cord, and Nicolas Thome · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Multi-modal factorized bilinear pooling with co-attention learning for visual question answering
Zhou Yu, Jun Yu, Jianping Fan, and Dacheng Tao · 2017
Earlier work this paper cites.
Don’t just assume; look and answer: Overcoming priors for visual question answering
Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi · 2018
Earlier work this paper cites.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang · 2018
Cited alongside, same era.
Women also snowboard: Overcoming bias in captioning models
Lisa Anne Hendricks, Kaylee Burns, Kate Saenko, Trevor Darrell, and Anna Rohrbach · 2018
Cited alongside, same era.
Jin-Hwa Kim, Jaehyun Jun, and Byoung-Tak Zhang · 2018
Cited alongside, same era.
Scn: switchable context network for semantic segmentation of rgb-d images
Di Lin, Ruimao Zhang, Yuanfeng Ji, Ping Li, and Hui Huang · 2018
Cited alongside, same era.
Overcoming language priors in visual question answering with adversarial regularization
Sainandan Ramakrishnan, Aishwarya Agrawal, and Stefan Lee · 2018
Cited alongside, same era.
Counterfactual samples synthesizing for robust visual question answering
Long Chen, Xin Yan, Jun Xiao, Hanwang Zhang, Shiliang Pu, and Yueting Zhuang · 2020
Later among the works it cites.
Mutant: A training paradigm for out-of-distribution generalization in visual question answering
Tejas Gokhale, Pratyay Banerjee, Chitta Baral, and Yezhou Yang · 2020
Later among the works it cites.
Compositional convolutional neural networks: A deep architecture with innate robustness to partial occlusion
Adam Kortylewski, Ju He, Qing Liu, and Alan L Yuille · 2020
Later among the works it cites.
Oscar: Object-semantics aligned pre-training for vision-language tasks
Xiujun Li, Xi Yin, Chunyuan Li, Pengchuan Zhang, Xiaowei Hu, Lei Zhang, Lijuan Wang, Houdong Hu, Li Dong, Furu Wei, et al · 2020
Later among the works it cites.
Gps-net: Graph property sensing network for scene graph generation
Xin Lin, Changxing Ding, Jinquan Zeng, and Dacheng Tao · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Beyond bilinear: Generalized multimodal factorized high-order pooling for visual question answering
Zhou Yu, Jun Yu, Chenchao Xiang, Jianping Fan, and Dacheng Tao · 2018
Cited alongside, same era.
Closure: Assessing systematic generalization of clevr models
Dzmitry Bahdanau, Harm de Vries, Timothy J O’Donnell, Shikhar Murty, Philippe Beaudoin, Yoshua Bengio, and Aaron Courville · 2019
Cited alongside, same era.
Murel: Multimodal relational reasoning for visual question answering
Rémi Cadène, Hedi Ben-younes, Matthieu Cord, and Nicolas Thome · 2019
Cited alongside, same era.
Rubi: Reducing unimodal biases for visual question answering
Remi Cadene, Corentin Dancette, Matthieu Cord, Devi Parikh, et al · 2019
Cited alongside, same era.
Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases
Christopher Clark, Mark Yatskar, and Luke Zettlemoyer · 2019
Cited alongside, same era.
Adaptive pyramid context network for semantic segmentation
Junjun He, Zhongying Deng, Lei Zhou, Yali Wang, and Yu Qiao · 2019
Cited alongside, same era.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Drew A Hudson and Christopher D Manning · 2019
Cited alongside, same era.
Scene graph generation with hierarchical context
Guanghui Ren, Lejian Ren, Yue Liao, Si Liu, Bo Li, Jizhong Han, and Shuicheng Yan · 2020
Later among the works it cites.
Robust object detection under occlusion with context-aware compositionalnets
Angtian Wang, Yihong Sun, Adam Kortylewski, and Alan L Yuille · 2020
Later among the works it cites.
Sketching image gist: Human-mimetic hierarchical scene graph generation
Wenbin Wang, Ruiping Wang, Shiguang Shan, and Xilin Chen · 2020
Later among the works it cites.
Cascade region proposal and global context for deep object detection
Qiaoyong Zhong, Chao Li, Yingying Zhang, Di Xie, Shicai Yang, and Shiliang Pu · 2020
Later among the works it cites.
Beyond question-based biases: Assessing multimodal shortcut learning in visual question answering
Corentin Dancette, Rémi Cadène, Damien Teney, and Matthieu Cord · 2021
Later among the works it cites.
Roses are red, violets are blue… but should vqa expect them to?
Corentin Kervadec, Grigory Antipov, Moez Baccouche, and Christian Wolf · 2021
Later among the works it cites.
How transferable are reasoning patterns in vqa?
Corentin Kervadec, Theo Jaunet, Grigory Antipov, Moez Baccouche, Romain Vuillemot, and Christian Wolf · 2021
Later among the works it cites.
Compositional convolutional neural networks: A robust and interpretable model for object recognition under occlusion
Adam Kortylewski, Qing Liu, Angtian Wang, Yihong Sun, and Alan Yuille · 2021
Later among the works it cites.
Counterfactual vqa: A cause-effect look at language bias
Yulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu, Xiansheng Hua, and Ji-Rong Wen · 2021
Later among the works it cites.
Counterfactual vqa: A cause-effect look at language bias
Yulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu, Xian-Sheng Hua, and Ji-Rong Wen · 2021
Later among the works it cites.
Visual question answering model based on graph neural network and contextual attention
Himanshu Sharma and Anand Singh Jalal · 2021
Later among the works it cites.
Simvqa: Exploring simulated environments for visual question answering
Paola Cascante-Bonilla, Hui Wu, Letao Wang, Rogerio Feris, and Vicente Ordonez · 2022
Closest in time.