Fetching the paper…
Reading the bibliography…
Visual dialog is a challenging vision-language task, where a dialog agent needs to answer a series of questions through reasoning on the image content and dialog history.
Visualbert: A simple and performant baseline for vision and language
Liunian Harold Li, Mark Yatskar, Da Yin, Cho-Jui Hsieh, and Kai-Wei Chang. 2019 · 1908
Earlier work this paper cites.
Domain adaptive training BERT for response selection
Taesun Whang, Dongyub Lee, Chanhee Lee, Kisu Yang, Dongsuk Oh, and Heuiseok Lim. 2019 · 1908
Earlier work this paper cites.
UNITER: learning universal image-text representations
Yen-Chun Chen, Linjie Li, Licheng Yu, Ahmed El Kholy, Faisal Ahmed, Zhe Gan, Yu Cheng, and Jingjing Liu. 2019 · 1909
Earlier work this paper cites.
Van-Quang Nguyen, Masanori Suganuma, and Takayuki Okatani. 2019 · 1911
Earlier work this paper cites.
Large-scale pretraining for visual dialog: A simple state-of-the-art baseline
Vishvak Murahari, Dhruv Batra, Devi Parikh, and Abhishek Das. 2019 · 1912
Earlier work this paper cites.
Dialgraph: Sparse graph learning networks for visual dialog
Gi-Cheon Kang, Junseok Park, Hwaran Lee, Byoung-Tak Zhang, and Jin-Hwa Kim. 2020 · 2004
Earlier work this paper cites.
Multi-view attention networks for visual dialog
Sungjin Park, Taesun Whang, Yeochan Yoon, and Hueiseok Lim. 2020 · 2004
Earlier work this paper cites.
Learning to rank: from pairwise approach to listwise approach
Zhe Cao, Tao Qin, Tie-Yan Liu, Ming-Feng Tsai, and Hang Li. 2007 · 2007
Earlier work this paper cites.
Listwise approach to learning to rank: theory and algorithm
Fen Xia, Tie-Yan Liu, Jue Wang, Wensheng Zhang, and Hang Li. 2008 · 2008
Earlier work this paper cites.
A general approximation framework for direct optimization of information retrieval measures
Tao Qin, Tie-Yan Liu, and Hang Li. 2010 · 2010
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V. Le. 2014 · 2014
Earlier work this paper cites.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier. 2014 · 2014
Earlier work this paper cites.
VQA: visual question answering
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Earlier work this paper cites.
Faster R-CNN: towards real-time object detection with region proposal networks
Shaoqing Ren, Kaiming He, Ross B. Girshick, and Jian Sun. 2015 · 2015
Earlier work this paper cites.
Lei Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. 2016 · 2016
Earlier work this paper cites.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu, Mike Schuster, Zhifeng Chen, Quoc V. Le, Mohammad Norouzi, Wolfgang Macherey, Maxim Krikun, Yuan Cao, Qin Gao, Klaus Macherey, Jeff Klingner, Apurva Shah, Melvin Johnson, Xiaobing Liu, Lukasz Kaiser, Stephan Gouws, Yoshikiyo Kato, Taku Kudo, Hideto Kazawa, Keith Stevens, George Kurian, Nishant Patil, Wei Wang, Cliff Young, Jason Smith, Jason Riesa, Alex Rudnick, Oriol Vinyals, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2016 · 2016
Cited alongside, same era.
Visual dialog
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José M. F. Moura, Devi Parikh, and Dhruv Batra. 2017 · 2017
Cited alongside, same era.
Learning to reason: End-to-end module networks for visual question answering
Ronghang Hu, Jacob Andreas, Marcus Rohrbach, Trevor Darrell, and Kate Saenko. 2017 · 2017
Cited alongside, same era.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Ranjay Krishna, Yuke Zhu, Oliver Groth, Justin Johnson, Kenji Hata, Joshua Kravitz, Stephanie Chen, Yannis Kalantidis, Li-Jia Li, David A. Shamma, Michael S. Bernstein, and Li Fei-Fei. 2017 · 2017
Cited alongside, same era.
Recursive visual attention in visual dialog
Yulei Niu, Hanwang Zhang, Manli Zhang, Jianhong Zhang, Zhiwu Lu, and Ji-Rong Wen. 2019 · 2019
Later among the works it cites.
A corpus for reasoning about natural language grounded in photographs
Alane Suhr, Stephanie Zhou, Ally Zhang, Iris Zhang, Huajun Bai, and Yoav Artzi. 2019 · 2019
Later among the works it cites.
Videobert: A joint model for video and language representation learning
Chen Sun, Austin Myers, Carl Vondrick, Kevin Murphy, and Cordelia Schmid. 2019 · 2019
Later among the works it cites.
LXMERT: learning cross-modality encoder representations from transformers
Hao Tan and Mohit Bansal. 2019 · 2019
Later among the works it cites.
Making history matter: History-advantage sequence training for visual dialog
Tianhao Yang, Zheng-Jun Zha, and Hanwang Zhang. 2019 · 2019
Later among the works it cites.
From recognition to cognition: Visual commonsense reasoning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Best of both worlds: Transferring knowledge from discriminative learning to a generative visual dialog model
Jiasen Lu, Anitha Kannan, Jianwei Yang, Devi Parikh, and Dhruv Batra. 2017 · 2017
Cited alongside, same era.
Visual reference resolution using attention memory for visual dialog
Paul Hongsuck Seo, Andreas M. Lehrmann, Bohyung Han, and Leonid Sigal. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Visual coreference resolution in visual dialog using neural module networks
Satwik Kottur, José M. F. Moura, Devi Parikh, Dhruv Batra, and Marcus Rohrbach. 2018 · 2018
Cited alongside, same era.
Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Piyush Sharma, Nan Ding, Sebastian Goodman, and Radu Soricut. 2018 · 2018
Cited alongside, same era.
Are you talking to me? reasoned visual dialog generation through adversarial learning
Qi Wu, Peng Wang, Chunhua Shen, Ian D. Reid, and Anton van den Hengel. 2018 · 2018
Cited alongside, same era.
Fusion of detected objects in text for visual question answering
Chris Alberti, Jeffrey Ling, Michael Collins, and David Reitter. 2019 · 2019
Cited alongside, same era.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Rowan Zellers, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Later among the works it cites.
Reasoning visual dialogs with structural and partial observations
Zilong Zheng, Wenguan Wang, Siyuan Qi, and Song-Chun Zhu. 2019 · 2019
Later among the works it cites.
History for visual dialog: Do we really need it?
Shubham Agarwal, Trung Bui, Joon-Young Lee, Ioannis Konstas, and Verena Rieser. 2020 · 2020
Closest in time.
Iterative context-aware graph inference for visual dialog
Dan Guo, Hui Wang, Hanwang Zhang, Zheng-Jun Zha, and Meng Wang. 2020 · 2020
Closest in time.
Dualvd: An adaptive dual encoding model for deep visual understanding in visual dialogue
Xiaoze Jiang, Jing Yu, Zengchang Qin, Yingying Zhuang, Xingxing Zhang, Yue Hu, and Qi Wu. 2020 · 2020
Closest in time.
Modality-balanced models for visual dialogue
Hyounghun Kim, Hao Tan, and Mohit Bansal. 2020 · 2020
Closest in time.
Unicoder-vl: A universal encoder for vision and language by cross-modal pre-training
Gen Li, Nan Duan, Yuejian Fang, Ming Gong, and Daxin Jiang. 2020 · 2020
Closest in time.
Two causal principles for improving visual dialog
Jiaxin Qi, Yulei Niu, Jianqiang Huang, and Hanwang Zhang. 2020 · 2020
Closest in time.
VL-BERT: pre-training of generic visual-linguistic representations
Weijie Su, Xizhou Zhu, Yue Cao, Bin Li, Lewei Lu, Furu Wei, and Jifeng Dai. 2020 · 2020
Closest in time.
Unified vision-language pre-training for image captioning and VQA
Luowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu, Jason J. Corso, and Jianfeng Gao. 2020 · 2020
Closest in time.
Dual attention networks for visual reference resolution in visual dialog
Gi-Cheon Kang, Jaeseo Lim, and Byoung-Tak Zhang. 2019 · 2033
Closest in time.
Factor graph attention
Idan Schwartz, Seunghak Yu, Tamir Hazan, and Alexander G. Schwing. 2019 · 2048
Closest in time.