Fetching the paper…
Reading the bibliography…
This paper presents a new model for visual dialog, Recurrent Dual Attention Network (ReDAN), using multi-step reasoning to answer a series of questions about an image.
Image-question-answer synergistic network for visual dialog
Dalu Guo, Chang Xu, and Dacheng Tao. 2019 · 1902
Earlier work this paper cites.
Dual attention networks for visual reference resolution in visual dialog
Gi-Cheon Kang, Jaeseo Lim, and Byoung-Tak Zhang. 2019 · 1902
Earlier work this paper cites.
Making history matter: Gold-critic sequence training for visual dialog
Tianhao Yang, Zheng-Jun Zha, and Hanwang Zhang. 2019 · 1902
Earlier work this paper cites.
Generative visual dialogue system via adaptive reasoning and weighted likelihood estimation
Heming Zhang, Shalini Ghosh, Larry Heck, Stephen Walsh, Junting Zhang, Jie Zhang, and C-C Jay Kuo. 2019 · 1902
Earlier work this paper cites.
Relation-aware graph attention network for visual question answering
Linjie Li, Zhe Gan, Yu Cheng, and Jingjing Liu. 2019 · 1903
Earlier work this paper cites.
Idan Schwartz, Seunghak Yu, Tamir Hazan, and Alexander Schwing. 2019 · 1904
Earlier work this paper cites.
Reasoning visual dialogs with structural and partial observations
Zilong Zheng, Wenguan Wang, Siyuan Qi, and Song-Chun Zhu. 2019 · 1904
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Kyunghyun Cho, Bart Van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Generative adversarial nets
Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Auto-encoding variational bayes
Diederik P Kingma and Max Welling. 2014 · 2014
Earlier work this paper cites.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Recurrent models of visual attention
Volodymyr Mnih, Nicolas Heess, Alex Graves, et al. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Earlier work this paper cites.
Vqa: Visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
From captions to visual concepts and back
Hao Fang, Saurabh Gupta, Forrest Iandola, Rupesh K Srivastava, Li Deng, Piotr Dollár, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John C Platt, et al. 2015 · 2015
Earlier work this paper cites.
Draw: A recurrent neural network for image generation
Karol Gregor, Ivo Danihelka, Alex Graves, Danilo Jimenez Rezende, and Daan Wierstra. 2015 · 2015
Cited alongside, same era.
Faster r-cnn: Towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross Girshick, and Jian Sun. 2015 · 2015
Cited alongside, same era.
Learning structured output representation using deep conditional generative models
Kihyuk Sohn, Honglak Lee, and Xinchen Yan. 2015 · 2015
Cited alongside, same era.
Show and tell: A neural image caption generator
Oriol Vinyals, Alexander Toshev, Samy Bengio, and Dumitru Erhan. 2015 · 2015
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
Neural module networks
Visual reference resolution using attention memory for visual dialog
Paul Hongsuck Seo, Andreas Lehrmann, Bohyung Han, and Leonid Sigal. 2017 · 2017
Later among the works it cites.
Reasonet: Learning to stop reading in machine comprehension
Yelong Shen, Po-Sen Huang, Jianfeng Gao, and Weizhu Chen. 2017 · 2017
Later among the works it cites.
End-to-end optimization of goal-driven and visually grounded dialogue systems
Florian Strub, Harm De Vries, Jeremie Mary, Bilal Piot, Aaron Courville, and Olivier Pietquin. 2017 · 2017
Later among the works it cites.
Tips and tricks for visual question answering: Learnings from the 2017 challenge
Damien Teney, Peter Anderson, Xiaodong He, and Anton van den Hengel. 2018 · 2017
Later among the works it cites.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Dan Klein. 2016 · 2016
Cited alongside, same era.
Multimodal compact bilinear pooling for visual question answering and visual grounding
Akira Fukui, Dong Huk Park, Daylen Yang, Anna Rohrbach, Trevor Darrell, and Marcus Rohrbach. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
The goldilocks principle: Reading children’s books with explicit memory representations
Felix Hill, Antoine Bordes, Sumit Chopra, and Jason Weston. 2016 · 2016
Cited alongside, same era.
Iterative alternating neural attention for machine reading
Alessandro Sordoni, Philip Bachman, Adam Trischler, and Yoshua Bengio. 2016 · 2016
Cited alongside, same era.
Stacked attention networks for image question answering
Zichao Yang, Xiaodong He, Jianfeng Gao, Li Deng, and Alex Smola. 2016 · 2016
Cited alongside, same era.
Evaluating visual conversational agents via cooperative human-ai games
Prithvijit Chattopadhyay, Deshraj Yadav, Viraj Prabhu, Arjun Chandrasekaran, Abhishek Das, Stefan Lee, Dhruv Batra, and Devi Parikh. 2017 · 2017
Cited alongside, same era.
Language-based image editing with recurrent attentive models
Jianbo Chen, Yelong Shen, Jianfeng Gao, Jingjing Liu, and Xiaodong Liu. 2018 · 2018
Later among the works it cites.
Quac: Question answering in context
Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen-tau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Neural approaches to conversational ai
Jianfeng Gao, Michel Galley, and Lihong Li. 2018 · 2018
Later among the works it cites.
Dialog-based interactive image retrieval
Xiaoxiao Guo, Hui Wu, Yu Cheng, Steven Rennie, and Rogerio Schmidt Feris. 2018 · 2018
Later among the works it cites.
Compositional attention networks for machine reasoning
Drew A Hudson and Christopher D Manning. 2018 · 2018
Later among the works it cites.
Two can play this game: visual dialog with discriminative question generation and answering
Unnat Jain, Svetlana Lazebnik, and Alexander G Schwing. 2018 · 2018
Later among the works it cites.
Visual coreference resolution in visual dialog using neural module networks
Satwik Kottur, José MF Moura, Devi Parikh, Dhruv Batra, and Marcus Rohrbach. 2018 · 2018
Later among the works it cites.
Answerer in questioner’s mind for goal-oriented visual dialogue
Sang-Woo Lee, Yu-Jung Heo, and Byoung-Tak Zhang. 2018 · 2018
Later among the works it cites.
Stochastic answer networks for machine reading comprehension
Xiaodong Liu, Yelong Shen, Kevin Duh, and Jianfeng Gao. 2018 · 2018
Later among the works it cites.
Flipdial: A generative model for two-way visual dialogue
Daniela Massiceti, N Siddharth, Puneet K Dokania, and Philip HS Torr. 2018 · 2018
Later among the works it cites.
Recursive visual attention in visual dialog
Yulei Niu, Hanwang Zhang, Manli Zhang, Jianhong Zhang, Zhiwu Lu, and Ji-Rong Wen. 2018 · 2018
Later among the works it cites.
Coqa: A conversational question answering challenge
Siva Reddy, Danqi Chen, and Christopher D Manning. 2018 · 2018
Later among the works it cites.
Ask no more: Deciding when to guess in referential visual dialogue
Ravi Shekhar, Tim Baumgartner, Aashish Venkatesh, Elia Bruni, Raffaella Bernardi, and Raquel Fernandez. 2018 · 2018
Later among the works it cites.
Visual reasoning with multi-hop feature modulation
Florian Strub, Mathieu Seurin, Ethan Perez, Harm de Vries, Philippe Preux, Aaron Courville, Olivier Pietquin, et al. 2018 · 2018
Later among the works it cites.
Are you talking to me? reasoned visual dialog generation through adversarial learning
Qi Wu, Peng Wang, Chunhua Shen, Ian Reid, and Anton van den Hengel. 2018 · 2018
Later among the works it cites.
Goal-oriented visual question generation via intermediate rewards
Junjie Zhang, Qi Wu, Chunhua Shen, Jian Zhang, Jianfeng Lu, and Anton Van Den Hengel. 2018 · 2018
Later among the works it cites.