Fetching the paper…
Reading the bibliography…
Dialog systems need to understand dynamic visual scenes in order to have conversations with users about the objects and events around them.
“Probabilistic methods in spoken–dialogue systems,”
Steve J Young, · 2000
Earlier work this paper cites.
“Juplter: a telephone-based conversational interface for weather information,”
Victor Zue, Stephanie Seneff, James R Glass, Joseph Polifroni, Christine Pao, Timothy J Hazen, and Lee Hetherington, · 2000
Earlier work this paper cites.
“Spoken dialogue technology: enabling the conversational user interface,”
Michael F McTear, · 2002
Earlier work this paper cites.
“Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition,”
Sergio Guadarrama, Niveda Krishnamoorthy, Girish Malkarnenkar, Subhashini Venugopalan, Raymond Mooney, Trevor Darrell, and Kate Saenko, · 2013
Earlier work this paper cites.
“Large-scale video classification with convolutional neural networks,”
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei, · 2014
Earlier work this paper cites.
“Adam: A method for stochastic optimization,”
Diederik Kingma and Jimmy Ba, · 2014
Earlier work this paper cites.
“A neural conversational model,”
Oriol Vinyals and Quoc Le, · 2015
Earlier work this paper cites.
Ryan Lowe, Nissan Pow, Iulian Serban, and Joelle Pineau, · 2015
Earlier work this paper cites.
“VQA: Visual Question Answering,”
Stanislaw Antol, Aishwarya Agrawal, Jiasen Lu, Margaret Mitchell, Dhruv Batra, C. Lawrence Zitnick, and Devi Parikh, · 2015
Earlier work this paper cites.
“Video paragraph captioning using hierarchical recurrent neural networks,”
Haonan Yu, Jiang Wang, Zhiheng Huang, Yi Yang, and Wei Xu, · 2015
Cited alongside, same era.
“Learning spatiotemporal features with 3d convolutional networks,”
Du Tran, Lubomir D. Bourdev, Rob Fergus, Lorenzo Torresani, and Manohar Paluri, · 2015
Cited alongside, same era.
“Yin and Yang: Balancing and answering binary visual questions,”
Peng Zhang, Yash Goyal, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh, · 2016
Cited alongside, same era.
Abhishek Das, Satwik Kottur, Khushi Gupta, Avi Singh, Deshraj Yadav, José M. F. Moura, Devi Parikh, and Dhruv Batra, · 2016
Cited alongside, same era.
“Hollywood in homes: Crowdsourcing data collection for activity understanding,”
Gunnar A. Sigurdsson, Gül Varol, Xiaolong Wang, Ivan Laptev, Ali Farhadi, and Abhinav Gupta, · 2016
Cited alongside, same era.
“Learning cooperative visual dialog agents with deep reinforcement learning,”
Abhishek Das, Satwik Kottur, José M.F. Moura, Stefan Lee, and Dhruv Batra, · 2017
Later among the works it cites.
“Attention-based multimodal fusion for video description,”
Chiori Hori, Takaaki Hori, Teng-Yok Lee, Ziming Zhang, Bret Harsham, John R. Hershey, Tim K. Marks, and Kazuhiko Sumi, · 2017
Later among the works it cites.
“Quo vadis, action recognition? a new model and the kinetics dataset,”
Joao Carreira and Andrew Zisserman, · 2017
Later among the works it cites.
“Attention-based multimodal fusion for video description,”
Chiori Hori, Takaaki Hori, Teng-Yok Lee, Ziming Zhang, Bret Harsham, John R Hershey, Tim K Marks, and Kazuhiko Sumi, · 2017
Later among the works it cites.
“Early and late integration of audio features for automatic video description,”
Chiori Hori, Takaaki Hori, Tim K Marks, and John R Hershey, · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Msr-vtt: A large video description dataset for bridging video and language,”
Jun Xu, Tao Mei, Ting Yao, and Yong Rui, · 2016
Cited alongside, same era.
“Soundnet: Learning sound representations from unlabeled video,”
Yusuf Aytar, Carl Vondrick, and Antonio Torralba, · 2016
Cited alongside, same era.
“Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering,”
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh, · 2017
Cited alongside, same era.
S. Hershey, S. Chaudhuri, D. P. W. Ellis, J. F. Gemmeke, A. Jansen, R. C. Moore, M. Plakal, D. Platt, R. A. Saurous, B. Seybold, M. Slaney, R. J. Weiss, and K. Wilson, · 2017
Later among the works it cites.
“The kinetics human action video dataset,”
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al., · 2017
Later among the works it cites.
“Audio visual scene-aware dialog (avsd) challenge at dstc7,”
Huda Alamri, Vincent Cartillier, Raphael Gontijo Lopes, Abhishek Das, Jue Wang, Irfan Essa, Dhruv Batra, Devi Parikh, Anoop Cherian, Tim K Marks, and Chiori Hori, · 2018
Closest in time.