Fetching the paper…
Reading the bibliography…
Due to the significant advancement of Natural Language Processing and Computer Vision-based models, Visual Question Answering (VQA) systems are becoming more intelligent and advanced.
Edge focusing
Fredrik Bergholm. 1987 · 1987
Earlier work this paper cites.
Deep inside convolutional networks: Visualising image classification models and saliency maps
Karen Simonyan, Andrea Vedaldi, and Andrew Zisserman. 2013 · 2013
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
Mateusz Malinowski and Mario Fritz. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering. In Proceedings of the IEEE international conference on computer vision . 2425–2433
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Visual turing test for computer vision systems
Donald Geman, Stuart Geman, Neil Hallonquist, and Laurent Younes. 2015 · 2015
Earlier work this paper cites.
U-net: Convolutional networks for biomedical image segmentation. In International Conference on Medical image computing and computer-assisted intervention . Springer, 234–241
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. 2015 · 2015
Earlier work this paper cites.
Examples are not enough, learn to criticize! criticism for interpretability
Been Kim, Rajiv Khanna, and Oluwasanmi O Koyejo. 2016 · 2016
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A Shamma, et al · 2016
Earlier work this paper cites.
" Why should i trust you?" Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining . 1135–1144
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2016 · 2016
Earlier work this paper cites.
Improved techniques for training gans
Tim Salimans, Ian Goodfellow, Wojciech Zaremba, Vicki Cheung, Alec Radford, and Xi Chen. 2016 · 2016
Earlier work this paper cites.
Stacked attention networks for image question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition . 21–29
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola. 2016 · 2016
Earlier work this paper cites.
Yin and yang: Balancing and answering binary visual questions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 5014–5022
Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2016 · 2016
Earlier work this paper cites.
Mutan: Multimodal tucker fusion for visual question answering. In Proceedings of the IEEE international conference on computer vision . 2612–2620
Hedi Ben-Younes, Rémi Cadene, Matthieu Cord, and Nicolas Thome. 2017 · 2017
Earlier work this paper cites.
Interpretability of deep learning models: A survey of results. In 2017 IEEE smartworld, ubiquitous intelligence & computing, advanced & trusted computed, scalable computing & communications, cloud & big data computing, Internet of people and smart city innovation (smartworld/SCALCOM/UIC/ATC/CBDcom/IOP/SCI) . IEEE, 1–6
Supriyo Chakraborty, Richard Tomsett, Ramya Raghavendra, Daniel Harborne, Moustafa Alzantot, Federico Cerutti, Mani Srivastava, Alun Preece, Simon Julier, Raghuveer M Rao, et al · 2017
Earlier work this paper cites.
Human attention in visual question answering: Do humans and deep networks look at the same regions?
Abhishek Das, Harsh Agrawal, Larry Zitnick, Devi Parikh, and Dhruv Batra. 2017 · 2017
Earlier work this paper cites.
Towards a rigorous science of interpretable machine learning
Finale Doshi-Velez and Been Kim. 2017 · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . 6904–6913
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017 · 2017
Earlier work this paper cites.
Survey of visual question answering: Datasets and techniques
Akshay Kumar Gupta. 2017 · 2017
Earlier work this paper cites.
Image-to-image translation with conditional adversarial networks. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1125–1134
Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A Efros. 2017 · 2017
Cited alongside, same era.
Explaining nonlinear classification decisions with deep taylor decomposition
Grégoire Montavon, Sebastian Lapuschkin, Alexander Binder, Wojciech Samek, and Klaus-Robert Müller. 2017 · 2017
Cited alongside, same era.
Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision . 618–626
Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2017 · 2017
Cited alongside, same era.
Opening the black box of deep neural networks via information
Ravid Shwartz-Ziv and Naftali Tishby. 2017 · 2017
Cited alongside, same era.
Visual question answering: A survey of methods and datasets
Interpretable machine learning: definitions, methods, and applications
W James Murdoch, Chandan Singh, Karl Kumbier, Reza Abbasi-Asl, and Bin Yu. 2019 · 2019
Later among the works it cites.
Towards explaining recommendations through local surrogate models. In Proceedings of the 34th ACM/SIGAPP Symposium on Applied Computing . 1671–1678
Caio Nóbrega and Leandro Marinho. 2019 · 2019
Later among the works it cites.
Question-conditioned counterfactual image generation for vqa
Jingjing Pan, Yash Goyal, and Stefan Lee. 2019 · 2019
Later among the works it cites.
Visual question answering using deep learning: A survey and performance analysis
Yash Srivastava, Vaishnav Murali, Shiv Ram Dubey, and Snehasis Mukherjee. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Qi Wu, Damien Teney, Peng Wang, Chunhua Shen, Anthony Dick, and Anton van den Hengel. 2017 · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition . 6077–6086
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Cited alongside, same era.
Explaining explanations: An overview of interpretability of machine learning. In 2018 IEEE 5th International Conference on data science and advanced analytics (DSAA) . IEEE, 80–89
Leilani H Gilpin, David Bau, Ben Z Yuan, Ayesha Bajwa, Michael Specter, and Lalana Kagal. 2018 · 2018
Cited alongside, same era.
Generating counterfactual explanations with natural language
Lisa Anne Hendricks, Ronghang Hu, Trevor Darrell, and Zeynep Akata. 2018 · 2018
Cited alongside, same era.
Tell-and-Answer: Towards Explainable Visual Question Answering using Attributes and Captions. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing . 1338–1346
Qing Li, Jianlong Fu, Dongfei Yu, Tao Mei, and Jiebo Luo. 2018 · 2018
Cited alongside, same era.
Mapping Instructions to Actions in 3D Environments with Visual Goal Prediction. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing . 2667–2678
Dipendra Misra, Andrew Bennett, Valts Blukis, Eyvind Niklasson, Max Shatkhin, and Yoav Artzi. 2018 · 2018
Cited alongside, same era.
Spectral Normalization for Generative Adversarial Networks. In International Conference on Learning Representations
Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. 2018 · 2018
Cited alongside, same era.
Text-adaptive generative adversarial networks: manipulating images with natural language. In Proceedings of the 32nd International Conference on Neural Information Processing Systems . 42–51
Seonghyeon Nam, Yunji Kim, and Seon Joo Kim. 2018 · 2018
Cited alongside, same era.
Interpretable visual question answering by visual grounding from attention supervision mining. In 2019 ieee winter conference on applications of computer vision (wacv) . IEEE, 349–357
Yundong Zhang, Juan Carlos Niebles, and Alvaro Soto. 2019 · 2019
Later among the works it cites.
Counterfactual samples synthesizing for robust visual question answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 10800–10809
Long Chen, Xin Yan, Jun Xiao, Hanwang Zhang, Shiliang Pu, and Yueting Zhuang. 2020 · 2020
Later among the works it cites.
Explaining data-driven decisions made by AI systems: the counterfactual approach
Carlos Fernández-Loría, Foster Provost, and Xintian Han. 2020 · 2020
Later among the works it cites.
ViCE: visual counterfactual explanations for machine learning models. In Proceedings of the 25th International Conference on Intelligent User Interfaces . 531–535
Oscar Gomez, Steffen Holter, Jun Yuan, and Enrico Bertini. 2020 · 2020
Later among the works it cites.
Generative adversarial networks
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2020 · 2020
Later among the works it cites.
Why Spectral Normalization Stabilizes GANs: Analysis and Improvements
Zinan Lin, Vyas Sekar, and Giulia Fanti. 2020 · 2020
Later among the works it cites.
Causal interpretability for machine learning-problems, methods and evaluation
Raha Moraffah, Mansooreh Karami, Ruocheng Guo, Adrienne Raglin, and Huan Liu. 2020 · 2020
Later among the works it cites.
Learning what makes a difference from counterfactual examples and gradient supervision. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part X 16 . Springer, 580–599
Damien Teney, Ehsan Abbasnedjad, and Anton van den Hengel. 2020 · 2020
Later among the works it cites.
Counterfactual explanations for machine learning: A review
Sahil Verma, John Dickerson, and Keegan Hines. 2020 · 2020
Later among the works it cites.
Interpretable CNNs for object classification
Quanshi Zhang, Xin Wang, Ying Nian Wu, Huilin Zhou, and Song-Chun Zhu. 2020 · 2020
Later among the works it cites.
Overcoming language priors with self-supervised learning for visual question answering
Xi Zhu, Zhendong Mao, Chunxiao Liu, Peng Zhang, Bin Wang, and Yongdong Zhang. 2020 · 2020
Later among the works it cites.
Counterfactual vqa: A cause-effect look at language bias. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . 12700–12710
Yulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu, Xian-Sheng Hua, and Ji-Rong Wen. 2021 · 2021
Later among the works it cites.
Image manipulation with natural language using Two-sided Attentive Conditional Generative Adversarial Network
Dawei Zhu, Aditya Mogadala, and Dietrich Klakow. 2021 · 2021
Later among the works it cites.