Fetching the paper…
Reading the bibliography…
In this paper, we aim to obtain improved attention for a visual question answering (VQA) task.
Equilibrium points in n-person games
Nash, J. F. et al. (1950) · 1950
Earlier work this paper cites.
On information and sufficiency
Kullback, S. and Leibler, R. A. (1951) · 1951
Earlier work this paper cites.
Gorillas in our midst: Sustained inattentional blindness for dynamic events
Simons, D. J. and Chabris, C. F. (1999) · 1999
Earlier work this paper cites.
Jensen-shannon divergence and hilbert space embedding
Fuglede, B. and Topsoe, F. (2004) · 2004
Earlier work this paper cites.
Learning to predict where humans look
Judd, T., Ehinger, K., Durand, F., and Torralba, A. (2009) · 2009
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
Malinowski, M. and Fritz, M. (2014) · 2014
Earlier work this paper cites.
Grounded compositional semantics for finding and describing images with sentences
Socher, R., Karpathy, A., Le, Q. V., Manning, C. D., and Ng, A. Y. (2014) · 2014
Earlier work this paper cites.
Deep domain confusion: Maximizing for domain invariance
Tzeng, E., Hoffman, J., Zhang, N., Saenko, K., and Darrell, T. (2014) · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C. L., and Parikh, D. (2015) · 2015
Earlier work this paper cites.
Mind’s eye: A recurrent visual representation for image caption generation
Chen, X. and Lawrence Zitnick, C. (2015) · 2015
Earlier work this paper cites.
From captions to visual concepts and back
Fang, H., Gupta, S., Iandola, F., Srivastava, R., Deng, L., Dollár, P., Gao, J., He, X., Mitchell, M., Platt, J., et al. (2015) · 2015
Earlier work this paper cites.
Visual turing test for computer vision systems
Geman, D., Geman, S., Hallonquist, N., and Younes, L. (2015) · 2015
Earlier work this paper cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A. and Fei-Fei, L. (2015) · 2015
Earlier work this paper cites.
Exploring models and data for image question answering
Ren, M., Kiros, R., and Zemel, R. (2015) · 2015
Earlier work this paper cites.
Show and tell: A neural image caption generator
Vinyals, O., Toshev, A., Bengio, S., and Erhan, D. (2015) · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., and Bengio, Y. (2015) · 2015
Earlier work this paper cites.
Human Attention in Visual Question Answering: Do Humans and Deep Networks Look at the Same Regions?
Das, A., Agrawal, H., Zitnick, C. L., Parikh, D., and Batra, D. (2016) · 2016
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Fukui, A., Park, D. H., Yang, D., Rohrbach, A., Darrell, T., and Rohrbach, M. (2016) · 2016
Cited alongside, same era.
Densecap: Fully convolutional localization networks for dense captioning
Johnson, J., Karpathy, A., and Fei-Fei, L. (2016) · 2016
Cited alongside, same era.
Visual question answering with question representation update (qru)
Li, R. and Jia, J. (2016) · 2016
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering
Lu, J., Yang, J., Batra, D., and Parikh, D. (2016) · 2016
Cited alongside, same era.
Mutan: Multimodal tucker fusion for visual question answering
Ben, younes, H., Cadene, R., Cord, M., and Thome, N. (2017) · 2017
Later among the works it cites.
Learning cooperative visual dialog agents with deep reinforcement learning
Das, A., Kottur, S., Moura, J. M., Lee, S., and Batra, D. (2017) · 2017
Later among the works it cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D. (2017) · 2017
Later among the works it cites.
Hadamard Product for Low-rank Bilinear Pooling
Kim, J.-H., On, K. W., Lim, W., Kim, J., Ha, J.-W., and Zhang, B.-T. (2017) · 2017
Later among the works it cites.
Least squares generative adversarial networks
Mao, X., Li, Q., Xie, H., Lau, R. Y., Wang, Z., and Paul Smolley, S. (2017) · 2017
Later among the works it cites.
Grad-cam: Visual explanations from deep networks via gradient-based localization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Noh, H., Hongsuck Seo, P., and Han, B. (2016) · 2016
Cited alongside, same era.
Where to look: Focus regions for visual question answering
Shih, K. J., Singh, S., and Hoiem, D. (2016) · 2016
Cited alongside, same era.
Deep coral: Correlation alignment for deep domain adaptation
Sun, B. and Saenko, K. (2016) · 2016
Cited alongside, same era.
Dynamic memory networks for visual and textual question answering
Xiong, C., Merity, S., and Socher, R. (2016) · 2016
Cited alongside, same era.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
Xu, H. and Saenko, K. (2016) · 2016
Cited alongside, same era.
Attribute2image: Conditional image generation from visual attributes
Yan, X., Yang, J., Sohn, K., and Lee, H. (2016) · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Yang, Z., He, X., Gao, J., Deng, L., and Smola, A. (2016) · 2016
Cited alongside, same era.
Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D. (2017) · 2017
Later among the works it cites.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P., He, X., Buehler, C., Teney, D., Johnson, M., Gould, S., and Zhang, L. (2018) · 2018
Later among the works it cites.
Deep attention neural tensor network for visual question answering
Bai, Y., Fu, J., Zhao, T., and Mei, T. (2018) · 2018
Later among the works it cites.
Deriving machine attention from human rationales
Bao, Y., Chang, S., Yu, M., and Barzilay, R. (2018) · 2018
Later among the works it cites.
Bilinear attention networks
Kim, J.-H., Jun, J., and Zhang, B.-T. (2018) · 2018
Later among the works it cites.
Differential attention for visual question answering
Patro, Badri, Namboodiri, and P, V. (2018) · 2018
Later among the works it cites.
Rise: Randomized input sampling for explanation of black-box models
Petsiuk, V., Das, A., and Saenko, K. (2018) · 2018
Later among the works it cites.
Learning to count objects in natural images for visual question answering
Zhang, Y., Hare, J., and Prügel-Bennett, A. (2018) · 2018
Later among the works it cites.
Attention is not explanation
Jain, S. and Wallace, B. C. (2019) · 2019
Closest in time.
U-cam: Visual explanation using uncertainty based class activation maps
Patro, B. N., Lunayach, M., Patel, S., and Namboodiri, V. P. (2019) · 2019
Closest in time.