Fetching the paper…
Reading the bibliography…
Visual dialog is a challenging vision-language task in which a series of questions visually grounded by a given image are answered.
The TREC-8 Question Answering Track Report
Voorhees, E. M.; et al. 1999 · 1999
Earlier work this paper cites.
Vd-bert: A unified vision and dialog transformer with bert
Wang, Y.; Joty, S.; Lyu, M. R.; King, I.; Xiong, C.; and Hoi, S. C. 2020 · 2004
Earlier work this paper cites.
Curriculum learning
Bengio, Y.; Louradour, J.; Collobert, R.; and Weston, J. 2009 · 2009
Earlier work this paper cites.
Microsoft coco: Common objects in context
Lin, T.-Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; and Zitnick, C. L. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J.; Socher, R.; and Manning, C. D. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D.; and Ba., J. 2015 · 2015
Earlier work this paper cites.
Faster r-cnn: Towards real-time object detection with region proposal networks
Ren, S.; He, K.; Girshick, R.; and Sun, J. 2015 · 2015
Earlier work this paper cites.
Neural module networks
Andreas, J.; Rohrbach, M.; Darrell, T.; and Klein, D. 2016 · 2016
Earlier work this paper cites.
Visual dialog
Das, A.; Kottur, S.; Gupta, K.; Singh, A.; Yadav, D.; Moura, J. M.; Parikh, D.; and Batra, D. 2017 · 2017
Earlier work this paper cites.
Visual genome: Connecting language and vision using crowdsourced dense image annotations
Krishna, R.; Zhu, Y.; Groth, O.; Johnson, J.; Hata, K.; Kravitz, J.; Chen, S.; Kalantidis, Y.; Li, L.-J.; Shamma, D. A.; et al. 2017 · 2017
Earlier work this paper cites.
Best of both worlds: Transferring knowledge from discriminative learning to a generative visual dialog model
Lu, J.; Kannan, A.; Yang, J.; Parikh, D.; and Batra, D. 2017 · 2017
Earlier work this paper cites.
Visual reference resolution using attention memory for visual dialog
Seo, P. H.; Lehrmann, A.; Han, B.; and Sigal, L. 2017 · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A. N.; Kaiser, Ł.; and Polosukhin, I. 2017 · 2017
Cited alongside, same era.
Bottom-up and top-down attention for image captioning and visual question answering
Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; and Zhang, L. 2018 · 2018
Cited alongside, same era.
Visual coreference resolution in visual dialog using neural module networks
Kottur, S.; Moura, J. M.; Parikh, D.; Batra, D.; and Rohrbach, M. 2018 · 2018
Cited alongside, same era.
Are you talking to me? reasoned visual dialog generation through adversarial learning
Wu, Q.; Wang, P.; Shen, C.; Reid, I.; and Van Den Hengel, A. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019 · 2019
Cited alongside, same era.
From recognition to cognition: Visual commonsense reasoning
Zellers, R.; Bisk, Y.; Farhadi, A.; and Choi, Y. 2019 · 2019
Later among the works it cites.
Reasoning visual dialogs with structural and partial observations
Zheng, Z.; Wang, W.; Qi, S.; and Zhu, S.-C. 2019 · 2019
Later among the works it cites.
History for Visual Dialog: Do we really need it?
Agarwal, S.; Bui, T.; Lee, J.-Y.; Konstas, I.; and Rieser, V. 2020 · 2020
Closest in time.
Iterative Context-Aware Graph Inference for Visual Dialog
Guo, D.; Wang, H.; Zhang, H.; Zha, Z.-J.; and Wang, M. 2020 · 2020
Closest in time.
DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual Dialogue
Jiang, X.; Yu, J.; Qin, Z.; Zhuang, Y.; Zhang, X.; Hu, Y.; and Wu, Q. 2020 · 2020
Closest in time.
Modality-balanced models for visual dialogue
Kim, H.; Tan, H.; and Bansal, M. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gan, Z.; Cheng, Y.; Kholy, A.; Li, L.; Liu, J.; and Gao, J. 2019 · 2019
Cited alongside, same era.
Image-question-answer synergistic network for visual dialog
Guo, D.; Xu, C.; and Tao, D. 2019 · 2019
Cited alongside, same era.
Gqa: A new dataset for real-world visual reasoning and compositional question answering
Hudson, D. A.; and Manning, C. D. 2019 · 2019
Cited alongside, same era.
Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Lu, J.; Batra, D.; Parikh, D.; and Lee, S. 2019 · 2019
Cited alongside, same era.
Recursive visual attention in visual dialog
Niu, Y.; Zhang, H.; Zhang, M.; Zhang, J.; Lu, Z.; and Wen, J.-R. 2019 · 2019
Cited alongside, same era.
PyTorch: An imperative style, high-performance deep learning library
Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. 2019 · 2019
Cited alongside, same era.
Making History Matter: History-Advantage Sequence Training for Visual Dialog
Yang, T.; Zha, Z.-J.; and Zhang, H. 2019 · 2019
Cited alongside, same era.
Closest in time.
Large-scale Pretraining for Visual Dialog: A Simple State-of-the-Art Baseline
Murahari, V.; Batra, D.; Parikh, D.; and Das, A. 2020 · 2020
Closest in time.
Efficient Attention Mechanism for Visual Dialog that can Handle All the Interactions between Multiple Inputs
Nguyen, V.-Q.; Suganuma, M.; and Okatani, T. 2020 · 2020
Closest in time.
Two Causal Principles for Improving Visual Dialog
Qi, J.; Niu, Y.; Huang, J.; and Zhang, H. 2020 · 2020
Closest in time.
Dual Attention Networks for Visual Reference Resolution in Visual Dialog
Kang, G.-C.; Lim, J.; and Zhang, B.-T. 2019 · 2033
Closest in time.
Factor graph attention
Schwartz, I.; Yu, S.; Hazan, T.; and Schwing, A. G. 2019 · 2048
Closest in time.