Fetching the paper…
Reading the bibliography…
Visual dialog is challenging since it needs to answer a series of coherent questions based on understanding the visual environment.
VD-BERT: A unified vision and dialog transformer with bert
Yue Wang, Shafiq Joty, Michael R Lyu, Irwin King, Caiming Xiong, and Steven CH Hoi. 2020 · 2004
Earlier work this paper cites.
Open domain dialogue generation with latent images
Ze Yang, Wei Wu, Huang Hu, Can Xu, and Zhoujun Li. 2020 · 2004
Earlier work this paper cites.
History for visual dialog: Do we really need it?
Shubham Agarwal, Trung Bui, Joon-Young Lee, Ioannis Konstas, and Verena Rieser. 2020 · 2005
Earlier work this paper cites.
Xiaoze Jiang, Jing Yu, Yajing Sun, Zengchang Qin, Zihao Zhu, Yue Hu, and Qi Wu. 2020c · 2007
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Are you talking to a machine? dataset and methods for multilingual image question
Haoyuan Gao, Junhua Mao, Jie Zhou, Zhiheng Huang, Lei Wang, and Wei Xu. 2015 · 2015
Earlier work this paper cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton. 2015 · 2015
Earlier work this paper cites.
SPICE: Semantic propositional image caption evaluation
Peter Anderson, Basura Fernando, Mark Johnson, and Stephen Gould. 2016 · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Hierarchical question-image co-attention for visual question answering
Jiasen Lu, Jianwei Yang, Dhruv Batra, and Devi Parikh. 2016 · 2016
Earlier work this paper cites.
Visual dialog
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José MF Moura, Devi Parikh, and Dhruv Batra. 2017 · 2017
Earlier work this paper cites.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017 · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. 2017 · 2017
Earlier work this paper cites.
Best of both worlds: Transferring knowledge from discriminative learning to a generative visual dialog model
Jiasen Lu, Anitha Kannan, Jianwei Yang, Devi Parikh, and Dhruv Batra. 2017 · 2017
Cited alongside, same era.
Visual reference resolution using attention memory for visual dialog
Paul Hongsuck Seo, Andreas Lehrmann, Bohyung Han, and Leonid Sigal. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Peter Anderson, Xiaodong He, Chris Buehler, Damien Teney, Mark Johnson, Stephen Gould, and Lei Zhang. 2018 · 2018
Cited alongside, same era.
Visual coreference resolution in visual dialog using neural module networks
Satwik Kottur, José M. F. Moura, Devi Parikh, Dhruv Batra, and Marcus Rohrbach. 2018 · 2018
Cited alongside, same era.
Meshed-memory transformer for image captioning
Marcella Cornia, Matteo Stefanini, Lorenzo Baraldi, and Rita Cucchiara. 2020 · 2020
Later among the works it cites.
Iterative context-aware graph inference for visual dialog
Dan Guo, Hui Wang, Hanwang Zhang, Zheng-Jun Zha, and Meng Wang. 2020 · 2020
Later among the works it cites.
Aligned dual channel graph convolutional network for visual question answering
Qingbao Huang, Jielong Wei, Yi Cai, Changmeng Zheng, Junying Chen, Ho-fung Leung, and Qing Li. 2020 · 2020
Later among the works it cites.
Large-scale pretraining for visual dialog: A simple state-of-the-art baseline
Vishvak Murahari, Dhruv Batra, Devi Parikh, and Abhishek Das. 2020 · 2020
Later among the works it cites.
Efficient attention mechanism for visual dialog that can handle all the interactions between multiple inputs
Van-Quang Nguyen, Masanori Suganuma, and Takayuki Okatani. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Image chat: Engaging grounded conversations
Kurt Shuster, Samuel Humeau, Antoine Bordes, and Jason Weston. 2018 · 2018
Cited alongside, same era.
Are you talking to me? reasoned visual dialog generation through adversarial learning
Qi Wu, Peng Wang, Chunhua Shen, Ian Reid, and Anton van den Hengel. 2018 · 2018
Cited alongside, same era.
Multi-step reasoning via recurrent dual attention for visual dialog
Zhe Gan, Yu Cheng, Ahmed EI Kholy, Linjie Li, Jingjing Liu, and Jianfeng Gao. 2019 · 2019
Cited alongside, same era.
What goes into a word: generating image descriptions with top-down spatial knowledge
Mehdi Ghanimifard and Simon Dobnik. 2019 · 2019
Cited alongside, same era.
Relation-aware graph attention network for visual question answering
Linjie Li, Zhe Gan, Yu Cheng, and Jingjing Liu. 2019 · 2019
Cited alongside, same era.
Recursive visual attention in visual dialog
Yulei Niu, Hanwang Zhang, Manli Zhang, Jianhong Zhang, Zhiwu Lu, and Ji-Rong Wen. 2019 · 2019
Cited alongside, same era.
Reasoning visual dialogs with structural and partial observations
Zilong Zheng, Wenguan Wang, Siyuan Qi, and Song-Chun Zhu. 2019 · 2019
Cited alongside, same era.
Jiaxin Qi, Yulei Niu, Jianqiang Huang, and Hanwang Zhang. 2020 · 2020
Later among the works it cites.
Gog: Relation-aware graph-over-graph network for visual dialog
Feilong Chen, Xiuyi Chen, Fandong Meng, Peng Li, and Jie Zhou. 2021a · 2021
Closest in time.
Multimodal incremental transformer with visual grounding for visual dialogue generation
Feilong Chen, Fandong Meng, Xiuyi Chen, Peng Li, and Jie Zhou. 2021c · 2021
Closest in time.
Unsupervised knowledge selection for dialogue generation
Xiuyi Chen, Feilong Chen, Fandong Meng, Peng Li, and Jie Zhou. 2021e · 2021
Closest in time.
Maria: A visual experience powered conversational agent
Zujie Liang, Huang Hu, Can Xu, Chongyang Tao, Xiubo Geng, Yining Chen, Fan Liang, and Daxin Jiang. 2021 · 2021
Closest in time.
Factor graph attention
Idan Schwartz, Seunghak Yu, Tamir Hazan, and Alexander G Schwing. 2019 · 2048
Closest in time.
Show, attend and tell: Neural image caption generation with visual attention
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron Courville, Ruslan Salakhudinov, Rich Zemel, and Yoshua Bengio. 2015 · 2057
Closest in time.