Fetching the paper…
Reading the bibliography…
Recent breakthroughs in computer vision and natural language processing have spurred interest in challenging multi-modal tasks such as visual question-answering and visual dialogue.
Everingham, M., Van Gool, L., Williams, C.K., Winn, J., Zisserman, A.: The pascal visual object classes (voc) challenge. International journal of computer vision 88
2010
Earlier work this paper cites.
Nair, V., Hinton, G.E.: Rectified linear units improve restricted boltzmann machines. In: Proc. of ICML (2010)
2010
Earlier work this paper cites.
Krizhevsky, A., Sutskever, I., Hinton, G.E.: Imagenet classification with deep convolutional neural networks. In: Proc. of of NIPS (2012)
2012
Earlier work this paper cites.
Mller, H., Clough, P., Deselaers, T., Caputo, B.: ImageCLEF: Experimental Evaluation in Visual Information Retrieval. Springer (2012)
2012
Earlier work this paper cites.
Girshick, R., Donahue, J., Darrell, T., Malik, J.: Rich feature hierarchies for accurate object detection and semantic segmentation. In: Proc. of of CVPR (2014)
2014
Earlier work this paper cites.
Graves, A., Wayne, G., Danihelka, I.: Neural turing machines. arXiv preprint arXiv:1410.5401 (2014)
2014
Earlier work this paper cites.
Kazemzadeh, S., Ordonez, V., Matten, M., Berg, T.: Referitgame: Referring to objects in photographs of natural scenes. In: Proc. of EMNLP (2014)
2014
Earlier work this paper cites.
Kingma, D.P., Ba, J.: Adam: A method for stochastic optimization. In: Proc. of ICLR (2014)
2014
Earlier work this paper cites.
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: Microsoft coco: Common objects in context. In: Proc. of ECCV (2014)
2014
Earlier work this paper cites.
Weston, J., Chopra, S., Bordes, A.: Memory networks. arXiv preprint arXiv:1410.3916 (2014)
2014
Earlier work this paper cites.
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Lawrence Zitnick, C., Parikh, D.: Vqa: Visual question answering. In: Proc. of ICCV (2015)
2015
Earlier work this paper cites.
Bahdanau, D., Cho, K., Bengio, Y.: Neural machine translation by jointly learning to align and translate. In: Proc. of ICLR (2015)
2015
Earlier work this paper cites.
Chung, J., Gulcehre, C., Cho, K., Bengio, Y.: Empirical evaluation of gated recurrent neural networks on sequence modeling. In: Proc. of ICML (2015)
2015
Earlier work this paper cites.
Ioffe, S., Szegedy, C.: Batch normalization: Accelerating deep network training by reducing internal covariate shift. In: Proc. of ICML (2015)
2015
Earlier work this paper cites.
Long, J., Shelhamer, E., Darrell, T.: Fully convolutional networks for semantic segmentation. In: Proc. of CVPR (2015)
2015
Earlier work this paper cites.
Luong, M.T., Pham, H., Manning, C.D.: Effective approaches to attention-based neural machine translation. In: Proc. of EMNLP (2015)
2015
Earlier work this paper cites.
Malinowski, M., Rohrbach, M., Fritz, M.: Ask your neurons: A neural-based approach to answering questions about images. In: Proc. of ICCV (2015)
2015
Earlier work this paper cites.
Russakovsky, O., Deng, J., Su, H., Krause, J., Satheesh, S., Ma, S., Huang, Z., Karpathy, A., Khosla, A., Bernstein, M., et al.: Imagenet large scale visual recognition challenge. International Journal of Computer Vision 115
2015
Earlier work this paper cites.
Sukhbaatar, S., Weston, J., Fergus, R., et al.: End-to-end memory networks. In: Proc. of NIPS (2015)
2015
Cited alongside, same era.
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., Bengio, Y.: Show, attend and tell: Neural image caption generation with visual attention. In: Proc. of ICML (2015)
2015
Cited alongside, same era.
Ba, J.L., Kiros, J.R., Hinton, G.E.: Layer normalization. Deep Learning Symposium (NIPS) (2016)
2016
Cited alongside, same era.
Fukui, A., Park, D.H., Yang, D., Rohrbach, A., Darrell, T., Rohrbach, M.: Multimodal compact bilinear pooling for visual question answering and visual grounding. In: Proc. of EMNLP (2016)
2016
Cited alongside, same era.
Graves, A., Wayne, G., Reynolds, M., Harley, T., Danihelka, I., Grabska-Barwińska, A., Colmenarejo, S.G., Grefenstette, E., Ramalho, T., Agapiou, J., et al.: Hybrid computing using a neural network with dynamic external memory. Nature 538
De Vries, H., Strub, F., Chandar, S., Pietquin, O., Larochelle, H., Courville, A.: Guesswhat?! visual object discovery through multi-modal dialogue. In: Proc. of CVPR (2017)
2017
Later among the works it cites.
Delbrouck, J.B., Dupont, S.: Modulating and attending the source image during encoding improves multimodal translation. Visually-Grounded Interaction and Language Workshop (NIPS) (2017)
2017
Later among the works it cites.
Dumoulin, V., Shlens, J., Kudlur, M.: A Learned Representation For Artistic Style. In: Proc. of ICLR (2017)
2017
Later among the works it cites.
Johnson, J., Hariharan, B., van der Maaten, L., Fei-Fei, L., Zitnick, C.L., Girshick, R.: Clevr: A diagnostic dataset for compositional language and elementary visual reasoning. In: Proc. of CVPR (2017)
2017
Later among the works it cites.
Kafle, K., Kanan, C.: Visual question answering: Datasets, algorithms, and future challenges. Computer Vision and Image Understanding 163
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2016
Cited alongside, same era.
He, K., Zhang, X., Ren, S., Sun, J.: Deep residual learning for image recognition. In: Proc. of CVPR (2016)
2016
Cited alongside, same era.
Hu, R., Rohrbach, M., Darrell, T.: Segmentation from natural language expressions. In: Proc. of ECCV (2016)
2016
Cited alongside, same era.
Hu, R., Xu, H., Rohrbach, M., Feng, J., Saenko, K., Darrell, T.: Natural language object retrieval. In: Proc. of CVPR (2016)
2016
Cited alongside, same era.
Jabri, A., Joulin, A., van der Maaten, L.: Revisiting visual question answering baselines. In: Proc. of ECCV (2016)
2016
Cited alongside, same era.
Kim, J.H., Lee, S.W., Kwak, D., Heo, M.O., Kim, J., Ha, J.W., Zhang, B.T.: Multimodal residual learning for visual qa. In: Proc. of NIPS (2016)
2016
Cited alongside, same era.
Lu, J., Yang, J., Batra, D., Parikh, D.: Hierarchical question-image co-attention for visual question answering. In: Proc. of NIPS (2016)
2016
Cited alongside, same era.
Nagaraja, V.K., Morariu, V.I., Davis, L.S.: Modeling context between objects for referring expression understanding. In: Proc. of ECCV (2016)
2016
Cited alongside, same era.
2017
Later among the works it cites.
Kim, J.H., On, K.W., Lim, W., Kim, J., Ha, J.W., Zhang, B.T.: Hadamard Product for Low-rank Bilinear Pooling. In: Proc. of ICLR (2017)
2017
Later among the works it cites.
Luo, R., Shakhnarovich, G.: Comprehension-guided referring expressions. In: Proc. of CVPR (2017)
2017
Later among the works it cites.
Strub, F., De Vries, H., Mary, J., Piot, B., Courville, A., Pietquin, O.: End-to-end optimization of goal-driven and visually grounded dialogue systems harm de vries. In: Proc. of IJCAI (2017)
2017
Later among the works it cites.
de Vries, H., Strub, F., Mary, J., Larochelle, H., Pietquin, O., Courville, A.C.: Modulating early visual processing by language. In: Proc. of NIPS (2017)
2017
Later among the works it cites.
Zhu, Y., Zhang, S., Metaxas, D.: Reasoning about fine-grained attribute phrases using reference games. In: Visually-Grounded Interaction and Language Workshop (NIPS) (2017)
2017
Later among the works it cites.
Dumoulin, V., Perez, E., Schucher, N., Strub, F., Vries, H.d., Courville, A., Bengio, Y.: Feature-wise transformations. Distill (2018). https://doi.org/10.23915/distill.00011, https://distill.pub/2018/feature-wise-transformations
2018
Closest in time.
Hudson, D.A., Manning, C.D.: Compositional attention networks for machine reasoning. In: Proc. of ICL (2018)
2018
Closest in time.
Lee, S.W., Heo, Y.J., Zhang, B.T.: Answerer in questioner’s mind for goal-oriented visual dialogue. Visually-Grounded Interaction and Language Workshop (NIPS) (2018)
2018
Closest in time.
Perez, E., Strub, F., De Vries, H., Dumoulin, V., Courville, A.: Film: Visual reasoning with a general conditioning layer. In: Proc. of AAAI (2018)
2018
Closest in time.
Rupprecht, C., Laina, I., Navab, N., Hager, G.D., Tombari, F.: Guide me: Interacting with deep networks. In: Proc. of CVPR (2018)
2018
Closest in time.
Yang, L., Wang, Y., Xiong, X., Yang, J., Katsaggelos, A.K.: Efficient video object segmentation via network modulation. In: Proc. of CVPR (2018)
2018
Closest in time.
Yu, L., Lin, Z., Shen, X., Yang, J., Lu, X., Bansal, M., Berg, T.L.: Mattnet: Modular attention network for referring expression comprehension. In: Proc. of CVPR (2018)
2018
Closest in time.