Fetching the paper…
Reading the bibliography…
In this paper, we propose a method to obtain robust explanations for visual question answering(VQA) that correlate well with the answers.
A model of inexact reasoning in medicine
E. H. Shortliffe and B. G. Buchanan · 1975
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
N. de freitas, d
K. Barnard, P. Duygulu, and D. Forsyth · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
S. Banerjee and A. Lavie · 2005
Earlier work this paper cites.
Explainable artificial intelligence for training and tutoring
H. C. Lane, M. G. Core, M. Van Lent, S. Solomon, and D. Gomboc · 2005
Earlier work this paper cites.
Statistical comparisons of classifiers over multiple data sets
J. Demšar · 2006
Earlier work this paper cites.
Every picture tells a story: Generating sentences from images
A. Farhadi, M. Hejrati, M. A. Sadeghi, P. Young, C. Rashtchian, J. Hockenmaier, and D. Forsyth · 2010
Earlier work this paper cites.
Baby talk: Understanding and generating image descriptions
G. Kulkarni, V. Premraj, S. Dhar, S. Li, Y. Choi, A. C. Berg, and T. L. Berg · 2011
Earlier work this paper cites.
What makes paris look like paris?
C. Doersch, S. Singh, A. Gupta, J. Sivic, and A. Efros · 2012
Earlier work this paper cites.
How do you tell a blackbird from a crow?
T. Berg and P. N. Belhumeur · 2013
Earlier work this paper cites.
Generative adversarial nets
I. Goodfellow, J. Pouget-Abadie, M. Mirza, B. Xu, D. Warde-Farley, S. Ozair, A. Courville, and Y. Bengio · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
M. Malinowski and M. Fritz · 2014
Earlier work this paper cites.
Conditional generative adversarial nets
M. Mirza and S. Osindero · 2014
Earlier work this paper cites.
Grounded compositional semantics for finding and describing images with sentences
R. Socher, A. Karpathy, Q. V. Le, C. D. Manning, and A. Y. Ng · 2014
Earlier work this paper cites.
Visualizing and understanding convolutional networks
M. D. Zeiler and R. Fergus · 2014
Earlier work this paper cites.
Object detectors emerge in deep scene cnns
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
Mind’s eye: A recurrent visual representation for image caption generation
X. Chen and C. Lawrence Zitnick · 2015
Earlier work this paper cites.
On the relationship between visual attributes and convolutional networks
V. Escorcia, J. Carlos Niebles, and B. Ghanem · 2015
Cited alongside, same era.
From captions to visual concepts and back
H. Fang, S. Gupta, F. Iandola, R. Srivastava, L. Deng, P. Dollár, J. Gao, X. He, M. Mitchell, J. Platt, et al · 2015
Cited alongside, same era.
Are you talking to a machine? dataset and methods for multilingual image question
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Cited alongside, same era.
Deep visual-semantic alignments for generating image descriptions
A. Karpathy and L. Fei-Fei · 2015
Cited alongside, same era.
Unsupervised representation learning with deep convolutional generative adversarial networks
A. Radford, L. Metz, and S. Chintala · 2015
Cited alongside, same era.
Exploring models and data for image question answering
X-cnn: Cross-modal convolutional neural networks for sparse datasets
P. Veličković, D. Wang, N. D. Lane, and P. Liò · 2016
Later among the works it cites.
Diverse beam search: Decoding diverse solutions from neural sequence models
A. K. Vijayakumar, M. Cogswell, R. R. Selvaraju, Q. Sun, S. Lee, D. Crandall, and D. Batra · 2016
Later among the works it cites.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
H. Xu and K. Saenko · 2016
Later among the works it cites.
Attribute2image: Conditional image generation from visual attributes
X. Yan, J. Yang, K. Sohn, and H. Lee · 2016
Later among the works it cites.
Visual7w: Grounded question answering in images
Y. Zhu, O. Groth, M. Bernstein, and L. Fei-Fei · 2016
Later among the works it cites.
Cross-modal scene networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Ren, R. Kiros, and R. Zemel · 2015
Cited alongside, same era.
Cider: Consensus-based image description evaluation
R. Vedantam, L. Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
K. Xu, J. Ba, R. Kiros, K. Cho, A. Courville, R. Salakhudinov, R. Zemel, and Y. Bengio · 2015
Cited alongside, same era.
Visual madlibs: Fill in the blank description generation and question answering
L. Yu, E. Park, A. C. Berg, and T. L. Berg · 2015
Cited alongside, same era.
Janes v0. 4: Korpus slovenskih spletnih uporabniških vsebin
D. Fišer, T. Erjavec, and N. Ljubešić · 2016
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
A. Fukui, D. H. Park, D. Yang, A. Rohrbach, T. Darrell, and M. Rohrbach · 2016
Cited alongside, same era.
Y. Aytar, L. Castrejon, C. Vondrick, H. Pirsiavash, and A. Torralba · 2017
Later among the works it cites.
Learning cooperative visual dialog agents with deep reinforcement learning
A. Das, S. Kottur, J. M. Moura, S. Lee, and D. Batra · 2017
Later among the works it cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Y. Goyal, T. Khot, D. Summers-Stay, D. Batra, and D. Parikh · 2017
Later among the works it cites.
Long text generation via adversarial training with leaked information
J. Guo, S. Lu, H. Cai, W. Zhang, Y. Yu, and J. Wang · 2017
Later among the works it cites.
Toward controlled generation of text
Z. Hu, Z. Yang, X. Liang, R. Salakhutdinov, and E. P. Xing · 2017
Later among the works it cites.
Creativity: Generating diverse questions using variational autoencoders
U. Jain, Z. Zhang, and A. Schwing · 2017
Later among the works it cites.
Adversarial learning for neural dialogue generation
J. Li, W. Monroe, T. Shi, A. Ritter, and D. Jurafsky · 2017
Later among the works it cites.
Recurrent topic-transition gan for visual paragraph generation
X. Liang, Z. Hu, H. Zhang, C. Gan, and E. P. Xing · 2017
Later among the works it cites.
Seqgan: Sequence generative adversarial nets with policy gradient
L. Yu, W. Zhang, J. Wang, and Y. Yu · 2017
Later among the works it cites.
Adversarial feature matching for text generation
Y. Zhang, Z. Gan, K. Fan, Z. Chen, R. Henao, D. Shen, and L. Carin · 2017
Later among the works it cites.
Multimodal explanations: Justifying decisions and pointing to the evidence
D. Huk Park, L. Anne Hendricks, Z. Akata, A. Rohrbach, B. Schiele, T. Darrell, and M. Rohrbach · 2018
Later among the works it cites.
Differential attention for visual question answering
B. Patro and V. P. Namboodiri · 2018
Later among the works it cites.
Multimodal differential network for visual question generation
B. N. Patro, S. Kumar, V. K. Kurmi, and V. Namboodiri · 2018
Later among the works it cites.
Explanation vs attention: A two-player game to obtain attention for vqa
B. N. Patro, Anupriy, and V. P. Namboodiri · 2019
Later among the works it cites.
U-cam: Visual explanation using uncertainty based class activation maps
B. N. Patro, M. Lunayach, S. Patel, and V. P. Namboodiri · 2019
Later among the works it cites.