Fetching the paper…
Reading the bibliography…
A major challenge for video captioning is to combine audio and visual cues.
Dropout: a simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. 2014 · 1958
Earlier work this paper cites.
Multimodal deep learning
Jiquan Ngiam, Aditya Khosla, Mingyu Kim, Juhan Nam, Honglak Lee, and Andrew Y Ng. 2011 · 2011
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler. 2012 · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Scheduled sampling for sequence prediction with recurrent neural networks
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer. 2015 · 2015
Earlier work this paper cites.
Microsoft coco captions: Data collection and evaluation server
Xinlei Chen, Hao Fang, Tsung-Yi Lin, Ramakrishna Vedantam, Saurabh Gupta, Piotr Dollár, and C Lawrence Zitnick. 2015 · 2015
Earlier work this paper cites.
Multimodal affective dimension prediction using deep bidirectional long short-term memory recurrent neural networks
Lang He, Dongmei Jiang, Le Yang, Ercheng Pei, Peng Wu, and Hichem Sahli. 2015 · 2015
Earlier work this paper cites.
Translating videos to natural language using deep recurrent neural networks
Subhashini Venugopalan, Huijuan Xu, Jeff Donahue, Marcus Rohrbach, Raymond Mooney, and Kate Saenko. 2015b · 2015
Earlier work this paper cites.
Describing videos by exploiting temporal structure
Li Yao, Atousa Torabi, Kyunghyun Cho, Nicolas Ballas, Christopher Pal, Hugo Larochelle, and Aaron Courville. 2015 · 2015
Earlier work this paper cites.
Multi-modal conditional attention fusion for dimensional emotion prediction
Shizhe Chen and Qin Jin. 2016 · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Cited alongside, same era.
Describing videos using multi-modal fusion
Qin Jin, Jia Chen, Shizhe Chen, Yifan Xiong, and Alexander Hauptmann. 2016 · 2016
Cited alongside, same era.
Hierarchical recurrent neural encoder for video representation with application to captioning
Pingbo Pan, Zhongwen Xu, Yi Yang, Fei Wu, and Yueting Zhuang. 2016 · 2016
Cited alongside, same era.
Multimodal video description
Vasili Ramanishka, Abir Das, Dong Huk Park, Subhashini Venugopalan, Lisa Anne Hendricks, Marcus Rohrbach, and Kate Saenko. 2016 · 2016
Cited alongside, same era.
Frame-and segment-level features and candidate pool evaluation for video caption generation
Cnn architectures for large-scale audio classification
Shawn Hershey, Sourish Chaudhuri, Daniel PW Ellis, Jort F Gemmeke, Aren Jansen, R Channing Moore, Manoj Plakal, Devin Platt, Rif A Saurous, Bryan Seybold, et al. 2017 · 2017
Later among the works it cites.
Attention-based multimodal fusion for video description
Chiori Hori, Takaaki Hori, Teng-Yok Lee, Ziming Zhang, Bret Harsham, John R. Hershey, Tim K. Marks, and Kazuhiko Sumi. 2017 · 2017
Later among the works it cites.
Reinforced video captioning with entailment rewards
Ramakanth Pasunuru and Mohit Bansal. 2017 · 2017
Later among the works it cites.
A deep reinforced model for abstractive summarization
Romain Paulus, Caiming Xiong, and Richard Socher. 2017 · 2017
Later among the works it cites.
Weakly supervised dense video captioning
Zhiqiang Shen, Jianguo Li, Zhou Su, Minjun Li, Yurong Chen, Yu-Gang Jiang, and Xiangyang Xue. 2017 · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rakshith Shetty and Jorma Laaksonen. 2016 · 2016
Cited alongside, same era.
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. 2016 · 2016
Cited alongside, same era.
Video paragraph captioning using hierarchical recurrent neural networks
Haonan Yu, Jiang Wang, Zhiheng Huang, Yi Yang, and Wei Xu. 2016 · 2016
Cited alongside, same era.
Sequence to sequence-video to text
Subhashini Venugopalan, Marcus Rohrbach, Jeffrey Donahue, Raymond Mooney, Trevor Darrell, and Kate Saenko. 2015a
Cited in the paper.
Survey on audiovisual emotion recognition: databases, features, and data fusion strategies
Chung-Hsien Wu, Jen-Chun Lin, and Wen-Li Wei. 2014a
Cited in the paper.
Exploring inter-feature and inter-class relationships with deep neural networks for video classification
Zuxuan Wu, Yu-Gang Jiang, Jun Wang, Jian Pu, and Xiangyang Xue. 2014b
Cited in the paper.
Learning multimodal attention lstm networks for video captioning
Jun Xu, Ting Yao, Yongdong Zhang, and Tao Mei. 2017 · 2017
Later among the works it cites.
Deep multimodal representation learning from temporal data
Xitong Yang, Palghat Ramesh, Radha Chitta, Sriganesh Madhvanath, Edgar A. Bernal, and Jiebo Luo. 2017 · 2017
Later among the works it cites.
Video captioning via hierarchical reinforcement learning
Xin Wang, Wenhu Chen, Jiawei Wu, Yuan-Fang Wang, and William Yang Wang. 2018 · 2018
Closest in time.