Fetching the paper…
Reading the bibliography…
Visual dialog is a task of answering a series of inter-dependent questions given an input image, and often requires to resolve visual references among the questions.
Imagenet: A large-scale hierarchical image database
Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: · 2009
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., Zitnick, C.L.: · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D., Ba, J.: · 2014
Earlier work this paper cites.
Show and tell: A neural image caption generator
Vinyals, O., Toshev, A., Bengio, S., Erhan, D.: · 2015
Earlier work this paper cites.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Courville, A., Salakhutdinov, R., Zemel, R., Bengio, Y.: · 2015
Earlier work this paper cites.
VQA: Visual Question Answering
Antol, S., Agrawal, A., Lu, J., Mitchell, M., Batra, D., Zitnick, C.L., Parikh, D.: · 2015
Earlier work this paper cites.
Ask your neurons: A neural-based approach to answering questions about images
Malinowski, M., Rohrbach, M., Fritz, M.: · 2015
Earlier work this paper cites.
Uncovering temporal context for video question and answering
Zhu, L., Xu, Z., Yang, Y., Hauptmann, A.G.: · 2015
Earlier work this paper cites.
End-to-end memory networks
Sukhbaatar, S., Weston, J., Fergus, R., et al.: · 2015
Earlier work this paper cites.
Memory networks
Weston, J., Chopra, S., Bordes, A.: · 2015
Earlier work this paper cites.
Entity-centric coreference resolution with model stacking
Clark, K., Manning, C.D.: · 2015
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Simonyan, K., Zisserman, A.: · 2015
Earlier work this paper cites.
Text-guided attention model for image captioning
Mun, J., Cho, M., Han, B.: · 2016
Earlier work this paper cites.
Grounding of textual phrases in images by reconstruction
Rohrbach, A., Rohrbach, M., Hu, R., Darrell, T., Schiele, B.: · 2016
Earlier work this paper cites.
Generating images from captions with attention
Mansimov, E., Parisotto, E., Ba, J., Salakhutdinov, R.: · 2016
Cited alongside, same era.
Generative adversarial text to image synthesis
Reed, S., Akata, Z., Yan, X., Logeswaran, L., Schiele, B., Lee, H.: · 2016
Cited alongside, same era.
Image question answering using convolutional neural network with dynamic parameter prediction
Noh, H., Seo, P.H., Han, B.: · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Yang, Z., He, X., Gao, J., Deng, L., Smola, A.: · 2016
Cited alongside, same era.
Ask, attend and answer: Exploring question-guided spatial attention for visual question answering
Xu, H., Saenko, K.: · 2016
Cited alongside, same era.
Deep compositional question answering with neural module networks
Andreas, J., Rohrbach, M., Darrell, T., Klein, D.: · 2016
Progressive attention networks for visual attribute prediction
Seo, P.H., Lin, Z., Cohen, S., Shen, X., Han, B.: · 2016
Later among the works it cites.
Ask me anything: Dynamic memory networks for natural language processing
Kumar, A., Irsoy, O., Ondruska, P., Iyyer, M., Bradbury, J., Gulrajani, I., Zhong, V., Paulus, R., Socher, R.: · 2016
Later among the works it cites.
Dynamic memory networks for visual and textual question answering
Xiong, C., Merity, S., Socher, R.: · 2016
Later among the works it cites.
Key-value memory networks for directly reading documents
Miller, A., Fisch, A., Dodge, J., Karimi, A.H., Bordes, A., Weston, J.: · 2016
Later among the works it cites.
Deep reinforcement learning for mention-ranking coreference models
Clark, K., Manning, C.D.: · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Neural module networks
Andreas, J., Rohrbach, M., Darrell, T., Klein, D.: · 2016
Cited alongside, same era.
Hierarchical question-image co-attention for visual question answering
Lu, J., Yang, J., Batra, D., Parikh, D.: · 2016
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Fukui, A., Park, D.H., Yang, D., Rohrbach, A., Darrell, T., Rohrbach, M.: · 2016
Cited alongside, same era.
Training recurrent answering units with joint loss minimization for vqa
Noh, H., Han, B.: · 2016
Cited alongside, same era.
Yin and Yang: Balancing and answering binary visual questions
Zhang, P., Goyal, Y., Summers-Stay, D., Batra, D., Parikh, D.: · 2016
Cited alongside, same era.
MarioQA: Answering questions by watching gameplay videos
Mun, J., Seo, P.H., Jung, I., Han, B.: · 2016
Cited alongside, same era.
Clark, K., Manning, C.D.: · 2016
Later among the works it cites.
Unsupervised visual-linguistic reference resolution in instructional videos
Huang, D.A., Lim, J.J., Fei-Fei, L., Niebles, J.C.: · 2017
Closest in time.
Hadamard Product for Low-rank Bilinear Pooling
Kim, J.H., On, K.W., Lim, W., Kim, J., Ha, J.W., Zhang, B.T.: · 2017
Closest in time.
Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., Parikh, D.: · 2017
Closest in time.
Visual Dialog
Das, A., Kottur, S., Gupta, K., Singh, A., Yadav, D., Moura, J.M., Parikh, D., Batra, D.: · 2017
Closest in time.
Guesswhat?! visual object discovery through multi-modal dialogue
de Vries, H., Strub, F., Chandar, S., Pietquin, O., Larochelle, H., Courville, A.: · 2017
Closest in time.
Learning cooperative visual dialog agents with deep reinforcement learning
Das, A., Kottur, S., Moura, J.M., Lee, S., Batra, D.: · 2017
Closest in time.
End-to-end optimization of goal-driven and visually grounded dialogue systems
Strub, F., de Vries, H., Mary, J., Piot, B., Courville, A., Pietquin, O.: · 2017
Closest in time.