Fetching the paper…
Reading the bibliography…
We address the problem of Visual Question Answering (VQA), which requires joint image and language understanding to answer a question about a given photograph.
Natural language questions for the web of data
Yahya, M., Berberich, K., Elbassuoni, S., Ramanath, M., Tresp, V., Weikum, G.: · 2012
Earlier work this paper cites.
Indoor segmentation and support inference from rgbd images
Nathan Silberman, Derek Hoiem, P.K., Fergus, R.: · 2012
Earlier work this paper cites.
A multi-world approach to question answering about real-world scenes based on uncertain input
Malinowski, M., Fritz, M.: · 2014
Earlier work this paper cites.
Joint video and text parsing for understanding events and answering queries
Tu, K., Meng, M., Lee, M.W., Choe, T.E., Zhu, S.C.: · 2014
Earlier work this paper cites.
Increasing the bandwidth of crowdsourced visual question answering to better support blind users
Lasecki, W.S., Zhong, Y., Bigham, J.P.: · 2014
Earlier work this paper cites.
Long-term recurrent convolutional networks for visual recognition and description
Donahue, J., Hendricks, L.A., Guadarrama, S., Rohrbach, M., Venugopalan, S., Saenko, K., Darrell, T.: · 2014
Earlier work this paper cites.
Show and tell: A neural image caption generator
Vinyals, O., Toshev, A., Bengio, S., Erhan, D.: · 2014
Earlier work this paper cites.
Deep fragment embeddings for bidirectional image sentence mapping
Karpathy, A., Joulin, A., Li, F.F.F.: · 2014
Earlier work this paper cites.
From captions to visual concepts and back
Fang, H., Gupta, S., Iandola, F., Srivastava, R., Deng, L., Dollár, P., Gao, J., He, X., Mitchell, M., Platt, J., et al.: · 2014
Earlier work this paper cites.
Weston, J., Chopra, S., Bordes, A.: · 2014
Earlier work this paper cites.
Caffe: Convolutional architecture for fast feature embedding
Jia, Y., Shelhamer, E., Donahue, J., Karayev, S., Long, J., Girshick, R., Guadarrama, S., Darrell, T.: · 2014
Earlier work this paper cites.
Semantic parsing via paraphrasing
Berant, J., Liang, P.: · 2014
Cited alongside, same era.
Question answering with subgraph embeddings
Bordes, A., Chopra, S., Weston, J.: · 2014
Cited alongside, same era.
Graves, A., Wayne, G., Danihelka, I.: · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., Bengio, Y.: · 2014
Cited alongside, same era.
Microsoft coco: Common objects in context
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: · 2014
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Courville, A., Salakhutdinov, R., Zemel, R., Bengio, Y.: · 2015
Closest in time.
Describing videos by exploiting temporal structure
Yao, L., Torabi, A., Cho, K., Ballas, N., Pal, C., Larochelle, H., Courville, A.: · 2015
Closest in time.
Effective approaches to attention-based neural machine translation
Luong, M.T., Pham, H., Manning, C.D.: · 2015
Closest in time.
Teaching machines to read and comprehend
Hermann, K.M., Kocisky, T., Grefenstette, E., Espeholt, L., Kay, W., Suleyman, M., Blunsom, P.: · 2015
Closest in time.
Describing multimedia content using attention-based encoder–decoder networks
Cho, K., Courville, A., Bengio, Y.: · 2015
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Venugopalan, S., Xu, H., Donahue, J., Rohrbach, M., Mooney, R., Saenko, K.: · 2014
Cited alongside, same era.
VQA: visual question answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., Parikh, D.: · 2015
Cited alongside, same era.
Simple baseline for visual question answering
Zhou, B., Tian, Y., Sukhbaatar, S., Szlam, A., Fergus, R.: · 2015
Cited alongside, same era.
Ask your neurons: A neural-based approach to answering questions about images
Malinowski, M., Rohrbach, M., Fritz, M.: · 2015
Cited alongside, same era.
Exploring models and data for image question answering
Ren, M., Kiros, R., Zemel, R.S.: · 2015
Cited alongside, same era.
Sukhbaatar, S., Szlam, A., Weston, J., Fergus, R.: · 2015
Cited alongside, same era.
Zhu, Y., Groth, O., Bernstein, M., Fei-Fei, L.: · 2015
Closest in time.
Ask me anything: Free-form visual question answering based on knowledge from external sources
Wu, Q., Wang, P., Shen, C., Hengel, A.v.d., Dick, A.: · 2015
Closest in time.
Image question answering using convolutional neural network with dynamic parameter prediction
Noh, H., Seo, P.H., Han, B.: · 2015
Closest in time.
Going deeper with convolutions
Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., Rabinovich, A.: · 2015
Closest in time.
Where to look: Focus regions for visual question answering
Shih, K.J., Singh, S., Hoiem, D.: · 2015
Closest in time.