Fetching the paper…
Reading the bibliography…
Visual question answering (VQA) is an interesting learning setting for evaluating the abilities and shortcomings of current systems for image understanding.
Dropout: A simple way to prevent neural networks from overfitting
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., Salakhutdinov, R.: · 1958
Earlier work this paper cites.
Long short-term memory
Hochreiter, S., Schmidhuber, J.: · 1997
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Koehn, P.: · 2004
Earlier work this paper cites.
Re-evaluation the role of bleu in machine translation research
Callison-Burch, C., Osborne, M., Koehn, P.: · 2006
Earlier work this paper cites.
Dbpedia: A nucleus for a web of open data
Auer, S., Bizer, C., Kobilarov, G., Lehmann, J., Cyganiak, R., Ives, Z.: · 2007
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., Hinton, G.: · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., Dean, J.: · 2013
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
Malinowski, M., Fritz, M.: · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollar, P., Zitnick, C.: · 2014
Earlier work this paper cites.
CNN features off-the-shelf: an astounding baseline for recognition
Razavian, A.S., Azizpour, H., Sullivan, J., Carlsson, S.: · 2014
Earlier work this paper cites.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., Rabinovich, A.: · 2015
Earlier work this paper cites.
Visual Turing test for computer vision systems
Geman, D., Geman, S., Hallonquist, N., Younes, L.: · 2015
Cited alongside, same era.
VQA: Visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C., Parikh, D.: · 2015
Cited alongside, same era.
Exploring models and data for image question answering
Ren, M., Kiros, R., Zemel, R.: · 2015
Cited alongside, same era.
Visual madlibs: Fill in the blank image generation and question answering
Yu, L., Park, E., Berg, A., Berg, T.: · 2015
Cited alongside, same era.
Visual7w: Grounded question answering in images
Zhu, Y., Groth, O., Bernstein, M., Fei-Fei, L.: · 2015
Cited alongside, same era.
Ask your neurons: A neural-based approach to answering questions about images
Learning visual features from large weakly supervised data
Joulin, A., van der Maaten, L., Jabri, A., Vasilache, N.: · 2015
Later among the works it cites.
Deep residual learning for image recognition
He, K., Zhang, X., Ren, S., Sun, J.: · 2016
Closest in time.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R., Zhu, Y., Groth, O., Johnson, J., Hata, K., Kravitz, J., Chen, S., Kalanditis, Y., Li, L.J., Shamma, D., Bernstein, M., Fei-Fei, L.: · 2016
Closest in time.
Human attention in visual question answering: Do humans and deep networks look at the same regions?
Das, A., Agrawal, H., Zitnick, C.L., Parikh, D., Batra, D.: · 2016
Closest in time.
Image captioning and visual question answering based on attributes and their related external knowledge
Wu, Q., Shen, C., van den Hengel, A., Wang, P., Dick, A.: · 2016
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Malinowski, M., Rohrbach, M., Fritz, M.: · 2015
Cited alongside, same era.
Are you talking to a machine? Dataset and methods for multilingual image question answering
Gao, H., Mao, J., Zhou, J., Huang, Z., Wang, L., Xu, W.: · 2015
Cited alongside, same era.
Simple baseline for visual question answering
Zhou, B., Tian, Y., Sukhbataar, S., Szlam, A., Fergus, R.: · 2015
Cited alongside, same era.
Learning to answer questions from image using convolutional neural network
Ma, L., Lu, Z., Li, H.: · 2015
Cited alongside, same era.
Deep compositional question answering with neural module networks
Andreas, J., Rohrbach, M., Darrell, T., Klein, D.: · 2015
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Fukui, A., Huk Park, D., Yang, D., Rohrbach, A., Darrell, T., Rohrbach, M.: · 2016
Closest in time.
Hierarchical question-image co-attention for visual question answering
Lu, J., Yang, J., Batra, D., Parikh, D.: · 2016
Closest in time.
Where to look: Focus regions for visual question answering
Shih, K.J., Singh, S., Hoiem, D.: · 2016
Closest in time.
Training and investigating residual nets
Gross, S., Wilber, M.: · 2016
Closest in time.
Multimodal residual learning for visual QA
Kim, J., Lee, S., Kwak, D., Heo, M., Kim, J., Ha, J., Zhang, B.: · 2016
Closest in time.