Fetching the paper…
Reading the bibliography…
Prior work in visual dialog has focused on training deep neural models on VisDial in isolation.
A. Krizhevsky, I. Sutskever, and G. Hinton, “ImageNet Classification with Deep Convolutional Neural Networks,” in
2012
Earlier work this paper cites.
P. Young, A. Lai, M. Hodosh, and J. Hockenmaier, “From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,”
2014
Earlier work this paper cites.
S. Kazemzadeh, V. Ordonez, M. Matten, and T. L. Berg, “ReferItGame: Referring to Objects in Photographs of Natural Scenes,” in
2014
Earlier work this paper cites.
T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona, D. Ramanan, P. Dollár, and C. L. Zitnick, “Microsoft COCO: Common Objects in Context,” in
2014
Earlier work this paper cites.
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh, “VQA: Visual Question Answering,” in
2015
Earlier work this paper cites.
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler, “Aligning books and movies: Towards story-like visual explanations by watching movies and reading books,” in
2015
Earlier work this paper cites.
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” in
2015
Earlier work this paper cites.
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. Bernstein,
2015
Earlier work this paper cites.
S. Ren, K. He, R. Girshick, and J. Sun, “Faster R-CNN: Towards real-time object detection with region proposal networks,” in
2015
Earlier work this paper cites.
D. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization,” in
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep Residual Learning for Image Recognition,” in
2016
Earlier work this paper cites.
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. Moura, D. Parikh, and D. Batra, “Visual Dialog,” in
2017
Earlier work this paper cites.
H. de Vries, F. Strub, S. Chandar, O. Pietquin, H. Larochelle, and A. Courville, “GuessWhat?! visual object discovery through multi-modal dialogue,” in
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
A. Das, S. Kottur, J. M. Moura, S. Lee, and D. Batra, “Learning Cooperative Visual Dialog Agents with Deep Reinforcement Learning,” in
2017
Earlier work this paper cites.
J. Lu, A. Kannan, J. Yang, D. Parikh, and D. Batra, “Best of both worlds: Transferring knowledge from discriminative learning to a generative visual dialog model,” in
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” in
2017
Earlier work this paper cites.
R. Krishna, Y. Zhu, O. Groth, J. Johnson, K. Hata, J. Kravitz, S. Chen, Y. Kalantidis, L.-J. Li, D. A. Shamma,
2017
Earlier work this paper cites.
2017
Earlier work this paper cites.
P. Sharma, N. Ding, S. Goodman, and R. Soricut, “Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,” in
2018
Earlier work this paper cites.
D. Massiceti, N. Siddharth, P. K. Dokania, and P. H. Torr, “Flipdial: A generative model for two-way visual dialogue,” in
2018
Earlier work this paper cites.
Q. Wu, P. Wang, C. Shen, I. Reid, and A. van den Hengel, “Are you talking to me? reasoned visual dialog generation through adversarial learning,” in
2018
Cited alongside, same era.
U. Jain, S. Lazebnik, and A. G. Schwing, “Two can play this game: visual dialog with discriminative question generation and answering,” in
2018
Cited alongside, same era.
S. Kottur, J. M. Moura, D. Parikh, D. Batra, and M. Rohrbach, “Visual coreference resolution in visual dialog using neural module networks,” in
2018
Cited alongside, same era.
2018
Cited alongside, same era.
A. Radford, K. Narasimhan, T. Salimans, and I. Sutskever, “Improving language understanding with unsupervised learning,” 2018
K. Nguyen and H. Daumé III, “Help, anna! visual navigation with natural multimodal assistance via retrospective curiosity-encouraging imitation learning,” in
2019
Closest in time.
J. Thomason, M. Murray, M. Cakmak, and L. Zettlemoyer, “Vision-and-dialog navigation,” 2019
2019
Closest in time.
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in
2019
Closest in time.
2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2018
Cited alongside, same era.
2018
Cited alongside, same era.
K.-H. Lee, X. Chen, G. Hua, H. Hu, and X. He, “Stacked cross attention for image-text matching,” in
2018
Cited alongside, same era.
J. Lu, D. Batra, D. Parikh, and S. Lee, “ViLBERT: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,” in
2019
Cited alongside, same era.
S.-W. Lee, T. Gao, S. Yang, J. Yoo, and J.-W. Ha, “Large-scale answerer in questioner’s mind for visual dialog question generation,” in
2019
Cited alongside, same era.
Y. Niu, H. Zhang, M. Zhang, J. Zhang, Z. Lu, and J.-R. Wen, “Recursive visual attention in visual dialog,” in
2019
Cited alongside, same era.
Z. Zheng, W. Wang, S. Qi, and S.-C. Zhu, “Reasoning visual dialogs with structural and partial observations,” in
2019
Cited alongside, same era.
I. Schwartz, S. Yu, T. Hazan, and A. G. Schwing, “Factor graph attention,” in
2019
Cited alongside, same era.
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
H. Tan and M. Bansal, “LXMERT: Learning cross-modality encoder representations from transformers,”
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
R. Zellers, Y. Bisk, A. Farhadi, and Y. Choi, “From recognition to cognition: Visual commonsense reasoning,” in
2019
Closest in time.
A. Suhr, S. Zhou, A. Zhang, I. Zhang, H. Bai, and Y. Artzi, “A corpus for reasoning about natural language grounded in photographs,” in
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
2019
Closest in time.
X. Jiang, J. Yu, Z. Qin, Y. Zhuang, X. Zhang, Y. Hu, and Q. Wu, “DualVD: An adaptive dual encoding model for deep visual understanding in visual dialogue,” in
2020
Closest in time.
W. Hao, C. Li, X. Li, L. Carin, and J. Gao, “Towards learning a generic agent for vision-and-language navigation via pre-training,” in
2020
Closest in time.