Fetching the paper…
Reading the bibliography…
Visual attention in Visual Question Answering (VQA) targets at locating the right image regions regarding the answer prediction, offering a powerful technique to promote multi-modal understanding.
Focal visual-text attention for visual question answering
Junwei Liang, Lu Jiang, Liangliang Cao, Li-Jia Li, and Alexander Hauptmann. 2019 · 1908
Earlier work this paper cites.
MuRel: Multimodal relational reasoning for visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . IEEE, 1989–1998
Remi Cadene, Hedi Ben-younes, Matthieu Cord, and Nicolas Thome. 2019a · 1998
Earlier work this paper cites.
VideoQA: question answering on news video. In ACM Multimedia . ACM, 632–641
Hui Yang, Lekha Chaisorn, Yunlong Zhao, Shi-Yong Neo, and Tat-Seng Chua. 2003 · 2003
Earlier work this paper cites.
Learning phrase representations using RNN Encoder-Decoder for statistical machine translation. In Proceedings of the Conference on Empirical Methods in Natural Language Processing . ACL, 1724–1734
Kyunghyun Cho, Bart van Merrienboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation. In Proceedings of the Conference on Empirical Methods in Natural Language Processing . ACL, 1532–1543
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering. In Proceedings of IEEE International Conference on Computer Vision . IEEE, 2425–2433
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn:Towards real-time object detection with region proposal networks. In Proceedings of Advances in Neural Information Processing Systems . Curran Associates, Inc., 91–99
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Deep compositional question answering with neural module networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . IEEE, 39–48
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. 2016 · 2016
Earlier work this paper cites.
Human attention in visual question answering: Do humans and deep networks look at the same regions?. In Proceedings of the Conference on Empirical Methods in Natural Language Processing . ACL, 932–937
Abhishek Das, Harsh Agrawal, C. Lawrence Zitnick, Devi Parikh, and Dhruv Batra. 2016 · 2016
Earlier work this paper cites.
Multimodal compact bilinear pooling for visual question answering and visual grounding. In Proceedings of the Conference on Empirical Methods in Natural Language Processing . ACL, 457–468
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach. 2016 · 2016
Earlier work this paper cites.
Hierarchical question-image co-attention for visual question answering. In Proceedings of Advances in Neural Information Processing Systems . Curran Associates, Inc., 289–297
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2016 · 2016
Earlier work this paper cites.
Where to look: Focus regions for visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . IEEE, 4613–4621
Kevin J. Shih, Saurabh Singh, and Derek Hoiem. 2016 · 2016
Earlier work this paper cites.
Stacked attention networks for image question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . IEEE, 21–29
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola. 2016 · 2016
Earlier work this paper cites.
Yin and yang: Balancing and answering binary visual questions. In Proceedings of Advances in Neural Information Processing Systems . Curran Associates, Inc., 5014–5022
Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2016 · 2016
Earlier work this paper cites.
Yuke Zhu, Oliver Groth, Michael Bernstein, and Li Fei-Fei. 2016 · 2016
Earlier work this paper cites.
Making the V in VQA matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . IEEE, 6325–6334
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017 · 2017
Earlier work this paper cites.
Right for the right reasons: Training differentiable models by constraining their explanations. In Proceedings of the International Joint Conference on Artificial Intelligence . ijcai.org, 2662–2670
Andrew Slavin Ross, Michael C. Hughes, and Finale Doshi-Velez. 2017 · 2017
Earlier work this paper cites.
Grad-CAM: Visual explanations from deep networks via gradient-based localization. In Proceedings of IEEE International Conference on Computer Vision . IEEE, 618–626
Ramprasaath R. Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2017 · 2017
Earlier work this paper cites.
Visual question answering: A survey of methods and datasets
Qi Wu, Damien Teney, Peng Wang, Chunhua Shen, Anthony Dick, and Anton van den Hengel. 2017 · 2017
Cited alongside, same era.
Video question answering via gradually refined attention over appearance and motion. In ACM Multimedia . ACM, 1645–1653
Dejing Xu, Zhou Zhao, Jun Xiao, Fei Wu, Hanwang Zhang, Xiangnan He, and Yueting Zhuang. 2017 · 2017
Cited alongside, same era.
Video question answering via hierarchical dual-level attention network learning. In ACM Multimedia . ACM, 1050–1058
Zhou Zhao, Jinghao Lin, Xinghua Jiang, Deng Cai, Xiaofei He, and Yueting Zhuang. 2017 · 2017
Cited alongside, same era.
Watch what you just said: Image captioning with text-conditional attention. In ACM Multimedia (Thematic Workshops) . ACM, 305–313
Luowei Zhou, Chenliang Xu, Parker A. Koch, and Jason J. Corso. 2017 · 2017
Cited alongside, same era.
Don’t just assume; look and answer: Overcoming priors for visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . IEEE, 4971–4980
Erasing-based attention learning for visual question answering. In ACM Multimedia . ACM, 1175–1183
Fei Liu, Jing Liu, Richang Hong, and Hanqing Lu. 2019 · 2019
Later among the works it cites.
Taking a hint: Leveraging explanations to make vision and language models more grounded. In Proceedings of IEEE International Conference on Computer Vision . IEEE, 2591–2600
Ramprasaath R. Selvaraju, Stefan Lee, Yilin Shen, Hongxia Jin, Shalini Ghosh, Larry Heck, Dhruv Batra, and Devi Parikh. 2019 · 2019
Later among the works it cites.
Self-critical reasoning for robust visual question answering. In Proceedings of Advances in Neural Information Processing Systems . Curran Associates, Inc., 8601–8611
Jialin Wu and Raymond J. Mooney. 2019 · 2019
Later among the works it cites.
Multi-source multi-level attention networks for visual question answering
Dongfei Yu, Jianlong Fu, Xinmei Tian, and Tao Mei. 2019 · 2019
Later among the works it cites.
Interpretable visual question answering by visual grounding from attention supervision mining. In Proceedings of the IEEE Winter Conference on Applications of Computer Vision . IEEE, 349–357
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Aishwarya Agrawal, Dhruv Batra, Devi Parikh, and Aniruddha Kembhavi. 2018 · 2018
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . IEEE, 6077–6086
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Cited alongside, same era.
Scale-aware Fast R-CNN for pedestrian detection
Jianan Li, Xiaodan Liang, Shengmei Shen, Tingfa Xu, Jiashi Feng, and Shuicheng Yan. 2018 · 2018
Cited alongside, same era.
Learning visual question answering by bootstrapping hard attention. In Proceedings of European Conference on Computer Vision . Springer, 3–20
Mateusz Malinowski, Carl Doersch, Adam Santoro, and Peter Battaglia. 2018 · 2018
Cited alongside, same era.
Exploring human-Like attention supervision in visual question answering. In AAAI . AAAI Press, 7300–7307
Tingting Qiao, Jianfeng Dong, and Duanqing Xu. 2018 · 2018
Cited alongside, same era.
Overcoming language priors in visual question answering with adversarial regularization. In Proceedings of Advances in Neural Information Processing Systems . Curran Associates, Inc., 1548–1558
Sainandan Ramakrishnan, Aishwarya Agrawal, and Stefan Lee. 2018 · 2018
Cited alongside, same era.
Don’t take the easy way out: Ensemble based methods for avoiding known dataset biases. In Proceedings of the Conference on Empirical Methods in Natural Language Processing . ACL, 4069–4082
Christopher Clark, Mark Yatskar, and Luke Zettlemoyer. 2019 · 2019
Cited alongside, same era.
BTDP: Toward sparse fusion with block term decomposition pooling for visual question answering
Zhiwei Fang, Jing Liu, Xueliang Liu, Qu Tang, Yong Li, and Hanqing Lu. 2019 · 2019
Cited alongside, same era.
Yundong Zhang, Juan Carlos Niebles, and Alvaro Soto. 2019 · 2019
Later among the works it cites.
Counterfactual samples synthesizing for robust visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . IEEE, 10800–10809
Long Chen, Xin Yan, Jun Xiao, Hanwang Zhang, Shiliang Pu, and Yueting Zhuang. 2020 · 2020
Later among the works it cites.
Recurrent attention network with reinforced generator for visual dialog
Hehe Fan, Linchao Zhu, Yi Yang, and Fei Wu. 2020 · 2020
Later among the works it cites.
Removing bias in multi-modal classifiers: regularization by maximizing functional entropies. In Proceedings of Advances in Neural Information Processing Systems . Curran Associates, Inc., 3197–3208
Itai Gat, Idan Schwartz, Alexander G. Schwing, and Tamir Hazan. 2020 · 2020
Later among the works it cites.
Loss-rescaling VQA: Revisiting language prior problem from a class-imbalance view
Yangyang Guo, Liqiang Nie, Zhiyong Cheng, and Qi Tian. 2020 · 2020
Later among the works it cites.
spaCy: Industrial-strength aatural language processing in python
Matthew Honnibal, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. 2020 · 2020
Later among the works it cites.
Overcoming language priors in VQA via decomposed linguistic representations. In AAAI . AAAI Press, 11181–11188
Chenchen Jing, Yuwei Wu, Xiaoxun Zhang, Yunde Jia, and Qi Wu. 2020 · 2020
Later among the works it cites.
Learning to contrast the counterfactual samples for robust visual question answering. In Proceedings of the Conference on Empirical Methods in Natural Language Processing . ACL, 3285–3292
Zujie Liang, Weitao Jiang, Haifeng Hu, and Jiaying Zhu. 2020 · 2020
Later among the works it cites.
Explanation vs attention: A two-player game to obtain attention for VQA. In AAAI . AAAI Press, 11848–11855
Badri N. Patro, Anupriy, and Vinay P. Namboodiri. 2020 · 2020
Later among the works it cites.
A negative case analysis of visual grounding methods for VQA. In Proceedings of the Conference on Empirical Methods in Natural Language Processing . ACL, 8172–8181
Robik Shrestha, Kushal Kafle, and Christopher Kanan. 2020 · 2020
Later among the works it cites.
On the value of out-of-distribution testing: An example of Goodhart’s law. In Proceedings of Advances in Neural Information Processing Systems . Curran Associates, Inc., 407–417
Damien Teney, Ehsan Abbasnejad, Kushal Kafle, Robik Shrestha, Christopher Kanan, and Anton van den Hengel. 2020 · 2020
Later among the works it cites.
Overcoming language priors with self-supervised learning for visual question answering. In Proceedings of the International Joint Conference on Artificial Intelligence . AAAI Press, 1083–1089
Xi Zhu, Zhendong Mao, Chunxiao Liu, Peng Zhang, Bin Wang, and Yongdong Zhang. 2020 · 2020
Later among the works it cites.
AdaVQA: Overcoming language priors with adapted margin cosine Loss. In Proceedings of the International Joint Conference on Artificial Intelligence . ijcai.org, 708–714
Yangyang Guo, Liqiang Nie, Zhiyong Cheng, Feng Ji, Ji Zhang, and Alberto Del Bimbo. 2021 · 2021
Closest in time.