Fetching the paper…
Reading the bibliography…
Video captioning is an advanced multi-modal task which aims to describe a video clip using a natural language sentence.
Ronald J. Williams and David Zipser, ‘A learning algorithm for continually running fully recurrent neural networks’, Neural Computation
1989
Earlier work this paper cites.
Sepp Hochreiter and Jürgen Schmidhuber, ‘Long short-term memory’, Neural Computation
1997
Earlier work this paper cites.
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu, ‘Bleu: a method for automatic evaluation of machine translation’, in ACL
2002
Earlier work this paper cites.
Chin-Yew Lin, ‘ROUGE: A package for automatic evaluation of summaries’, in Text Summarization Branches Out
2004
Earlier work this paper cites.
Satanjeev Banerjee and Alon Lavie, ‘METEOR: an automatic metric for MT evaluation with improved correlation with human judgments’, in Proceedings of the Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization@ACL
2005
Earlier work this paper cites.
Yoshua Bengio, Jérôme Louradour, Ronan Collobert, and Jason Weston, ‘Curriculum learning’, in Proceedings of the 26th Annual International Conference on Machine Learning
2009
Earlier work this paper cites.
Sergio Guadarrama, Niveda Krishnamoorthy, Girish Malkarnenkar, Subhashini Venugopalan, Raymond J. Mooney, Trevor Darrell, and Kate Saenko, ‘Youtube2text: Recognizing and describing arbitrary activities using semantic hierarchies and zero-shot recognition’, in ICCV
2013
Earlier work this paper cites.
Kyunghyun Cho, Bart Van Merrienboer, Dzmitry Bahdanau, and Yoshua Bengio, ‘On the properties of neural machine translation: Encoder-decoder approaches’, arXiv: Computation and Language
2014
Earlier work this paper cites.
Kyunghyun Cho, Bart van Merrienboer, Çaglar Gülçehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio, ‘Learning phrase representations using RNN encoder-decoder for statistical machine translation’, in EMNLP
2014
Earlier work this paper cites.
Junyoung Chung, Caglar Gulcehre, Kyunghyun Cho, and Yoshua Bengio, ‘Empirical evaluation of gated recurrent neural networks on sequence modeling’, arXiv: Neural and Evolutionary Computing
2014
Earlier work this paper cites.
Nitish Srivastava, Geoffrey E. Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov, ‘Dropout: a simple way to prevent neural networks from overfitting’, J. Mach. Learn. Res
2014
Earlier work this paper cites.
Wojciech Zaremba, Ilya Sutskever, and Oriol Vinyals, ‘Recurrent neural network regularization’, CoRR
2014
Earlier work this paper cites.
Jimmy Ba, Volodymyr Mnih, and Koray Kavukcuoglu, ‘Multiple object recognition with visual attention’, in ICLR
2015
Earlier work this paper cites.
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio, ‘Neural machine translation by jointly learning to align and translate’, in ICLR
2015
Earlier work this paper cites.
Samy Bengio, Oriol Vinyals, Navdeep Jaitly, and Noam Shazeer, ‘Scheduled sampling for sequence prediction with recurrent neural networks’, in NeurIPS
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
Sergey Ioffe and Christian Szegedy, ‘Batch normalization: Accelerating deep network training by reducing internal covariate shift’, in ICML
2015
Cited alongside, same era.
Taesup Moon, Heeyoul Choi, Hoshik Lee, and Inchul Song, ‘RNNDROP: A novel dropout for RNNS in ASR’, in IEEE Workshop on Automatic Speech Recognition and Understanding
2015
Cited alongside, same era.
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott E. Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich, ‘Going deeper with convolutions’, in CVPR
2015
Cited alongside, same era.
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh, ‘Cider: Consensus-based image description evaluation’, in CVPR
2015
Cited alongside, same era.
Subhashini Venugopalan, Huijuan Xu, Jeff Donahue, Marcus Rohrbach, Raymond Mooney, and Kate Saenko, ‘Translating videos to natural language using deep recurrent neural networks’, in NAACL
Chiori Hori, Takaaki Hori, Teng-Yok Lee, Ziming Zhang, Bret Harsham, John R. Hershey, Tim K. Marks, and Kazuhiko Sumi, ‘Attention-based multimodal fusion for video description’, in ICCV
2017
Later among the works it cites.
Ramakanth Pasunuru and Mohit Bansal, ‘Multi-task video captioning with video and entailment generation’, in ACL
2017
Later among the works it cites.
Ramakanth Pasunuru and Mohit Bansal, ‘Reinforced video captioning with entailment rewards’, in EMNLP
2017
Later among the works it cites.
Steven J. Rennie, Etienne Marcheret, Youssef Mroueh, Jerret Ross, and Vaibhava Goel, ‘Self-critical sequence training for image captioning’, in CVPR
2017
Later among the works it cites.
Saining Xie, Ross B. Girshick, Piotr Dollár, Zhuowen Tu, and Kaiming He, ‘Aggregated residual transformations for deep neural networks’, in CVPR
2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
Kelvin Xu, Jimmy Ba, Ryan Kiros, Kyunghyun Cho, Aaron C. Courville, Ruslan Salakhutdinov, Richard S. Zemel, and Yoshua Bengio, ‘Show, attend and tell: Neural image caption generation with visual attention’, in ICML
2015
Cited alongside, same era.
Li Yao, Atousa Torabi, Kyunghyun Cho, Nicolas Ballas, Christopher J. Pal, Hugo Larochelle, and Aaron C. Courville, ‘Describing videos by exploiting temporal structure’, in ICCV
2015
Cited alongside, same era.
Jimmy Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton, ‘Layer normalization’, ArXiv
2016
Cited alongside, same era.
Yarin Gal and Zoubin Ghahramani, ‘Dropout as a bayesian approximation: Representing model uncertainty in deep learning’, in ICML
2016
Cited alongside, same era.
Yarin Gal and Zoubin Ghahramani, ‘A theoretically grounded application of dropout in recurrent neural networks’, in NeurIPS
2016
Cited alongside, same era.
Anirudh Goyal, Alex Lamb, Ying Zhang, Saizheng Zhang, Aaron C. Courville, and Yoshua Bengio, ‘Professor forcing: A new algorithm for training recurrent networks’, in NeurIPS
2016
Cited alongside, same era.
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, ‘Deep residual learning for image recognition’, in CVPR
2016
Cited alongside, same era.
Later among the works it cites.
Chih-Yao Ma, Asim Kadav, Iain Melvin, Zsolt Kira, Ghassan AlRegib, and Hans Peter Graf, ‘Attend and interact: Higher-order object interactions for video understanding’, in CVPR
2018
Later among the works it cites.
Xin Wang, Wenhu Chen, Jiawei Wu, Yuan-Fang Wang, and William Yang Wang, ‘Video captioning via hierarchical reinforcement learning’, in CVPR
2018
Later among the works it cites.
Xin Wang, Yuan-Fang Wang, and William Yang Wang, ‘Watch, listen, and describe: Globally and locally aligned cross-modal attentions for video captioning’, in NAACL-HLT
2018
Later among the works it cites.
Chunlei Wu, Yiwei Wei, Xiaoliang Chu, Weichen Sun, Fei Su, and Leiquan Wang, ‘Hierarchical attention-based multimodal fusion for video captioning’, Neurocomputing
2018
Later among the works it cites.
Mohammadreza Zolfaghari, Kamaljeet Singh, and Thomas Brox, ‘ECO: efficient convolutional network for online video understanding’, in ECCV
2018
Later among the works it cites.
Nayyer Aafaq, Naveed Akhtar, Wei Liu, Syed Zulqarnain Gilani, and Ajmal Mian, ‘Spatio-temporal dynamics and semantic attribute enriched visual encoding for video captioning’, in CVPR
2019
Later among the works it cites.
2019
Later among the works it cites.
Wenjie Pei, Jiyuan Zhang, Xiangrong Wang, Lei Ke, Xiaoyong Shen, and Yu-Wing Tai, ‘Memory-attended recurrent network for video captioning’, in CVPR
2019
Later among the works it cites.
Xin Wang, Jiawei Wu, Da Zhang, Yu Su, and William Yang Wang, ‘Learning to compose topic-aware mixture of experts for zero-shot video captioning’, in AAAI
2019
Later among the works it cites.
Chen Wu, Xuancheng Ren, Fuli Luo, and Xu Sun, ‘A hierarchical reinforced sequence operation method for unsupervised text style transfer’, in ACL
2019
Later among the works it cites.