Fetching the paper…
Reading the bibliography…
We describe a very simple bag-of-words baseline for visual question answering.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Multimodal neural language models
R. Kiros, R. Salakhutdinov, and R. Zemel · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick · 2014
Earlier work this paper cites.
Deep captioning with multimodal recurrent neural networks (m-rnn)
J. Mao, W. Xu, Y. Yang, J. Wang, and A. Yuille · 2014
Earlier work this paper cites.
Going deeper with convolutions
C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke, and A. Rabinovich · 2014
Earlier work this paper cites.
Show and tell: A neural image caption generator
O. Vinyals, A. Toshev, S. Bengio, and D. Erhan · 2014
Earlier work this paper cites.
Learning deep features for scene recognition using places database
B. Zhou, A. Lapedriza, J. Xiao, A. Torralba, and A. Oliva · 2014
Earlier work this paper cites.
Deep compositional question answering with neural module networks
J. Andreas, M. Rohrbach, T. Darrell, and D. Klein · 2015
Cited alongside, same era.
Vqa: Visual question answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Cited alongside, same era.
Abc-cnn: An attention based convolutional neural network for visual question answering
K. Chen, J. Wang, L.-C. Chen, H. Gao, W. Xu, and R. Nevatia · 2015
Cited alongside, same era.
Exploring nearest neighbor approaches for image captioning
J. Devlin, S. Gupta, R. Girshick, M. Mitchell, and C. L. Zitnick · 2015
Cited alongside, same era.
Are you talking to a machine? dataset and methods for multilingual image question answering
Exploring models and data for image question answering
M. Ren, R. Kiros, and R. Zemel · 2015
Closest in time.
Where to look: Focus regions for visual question answering
K. J. Shih, S. Singh, and D. Hoiem · 2015
Closest in time.
Ask me anything: Free-form visual question answering based on knowledge from external sources
Q. Wu, P. Wang, C. Shen, A. v. d. Hengel, and A. Dick · 2015
Closest in time.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
H. Xu and K. Saenko · 2015
Closest in time.
Stacked attention networks for image question answering
Z. Yang, X. He, J. Gao, L. Deng, and A. Smola · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Gao, J. Mao, J. Zhou, Z. Huang, L. Wang, and W. Xu · 2015
Cited alongside, same era.
Compositional memory for visual question answering
A. Jiang, F. Wang, F. Porikli, and Y. Li · 2015
Cited alongside, same era.
Image question answering using convolutional neural network with dynamic parameter prediction
H. Noh, P. H. Seo, and B. Han · 2015
Cited alongside, same era.
Closest in time.
Learning deep features for discriminative localization
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba · 2015
Closest in time.