Fetching the paper…
Reading the bibliography…
Scene-aware dialog systems will be able to have conversations with users about the objects and events around them.
Translating videos to natural language using deep recurrent neural networks
S. Venugopalan, H. Xu, J. Donahue, M. Rohrbach, R. Mooney, and K. Saenko · 2014
Earlier work this paper cites.
VQA: Visual Question Answering
S. Antol, A. Agrawal, J. Lu, M. Mitchell, D. Batra, C. L. Zitnick, and D. Parikh · 2015
Earlier work this paper cites.
The ubuntu dialogue corpus: A large dataset for research in unstructured multi-turn dialogue systems
R. Lowe, N. Pow, I. Serban, and J. Pineau · 2015
Earlier work this paper cites.
A. Das, S. Kottur, K. Gupta, A. Singh, D. Yadav, J. M. F. Moura, D. Parikh, and D. Batra · 2016
Cited alongside, same era.
Hollywood in homes: Crowdsourcing data collection for activity understanding
G. A. Sigurdsson, G. Varol, X. Wang, I. Laptev, A. Farhadi, and A. Gupta · 2016
Cited alongside, same era.
Learning cooperative visual dialog agents with deep reinforcement learning
A. Das, S. Kottur, J. M. Moura, S. Lee, and D. Batra · 2017
Cited alongside, same era.
Guesswhat?! visual object discovery through multi-modal dialogue
H. De Vries, F. Strub, S. Chandar, O. Pietquin, H. Larochelle, and A. Courville · 2017
Later among the works it cites.
End-to-end conversation modeling track in DSTC6
C. Hori and T. Hori · 2017
Later among the works it cites.
Attention-based multimodal fusion for video description
C. Hori, T. Hori, T.-Y. Lee, Z. Zhang, B. Harsham, J. R. Hershey, T. K. Marks, and K. Sumi · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…